IC-632BLIP-2 succeeds on only 5 out of 100 advanced compositional vision-language tasks

Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny

SourceMiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

The paper evaluates BLIP-2 on four advanced vision-language tasks (meme interpretation, recipe generation, advertisement creation, poem composition) with 25 images each, assessed by human evaluators. BLIP-2 succeeded on 0/25 memes, 4/25 recipes, 1/25 ads, and 0/25 poems, totalling 5/100. The paper attributes this to BLIP-2's use of FLAN-T5-XXL as its language model, which is less capable than the Vicuna LLM used in MiniGPT-4, and notes that even after fine-tuning BLIP-2 on the authors' second-stage data it still generates short responses and fails to generalise to advanced tasks.

Evidence
observational
Key metric
0/25 meme, 4/25 recipes, 1/25 ads, 0/25 poem, 5/100 avg
Model
BLIP-2
Concepts
Failure mode
Related findings
IC-633, IC-634, IC-635
Extraction
automatic-extraction