SourceMiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
The paper evaluates BLIP-2 on four advanced vision-language tasks (meme interpretation, recipe generation, advertisement creation, poem composition) with 25 images each, assessed by human evaluators. BLIP-2 succeeded on 0/25 memes, 4/25 recipes, 1/25 ads, and 0/25 poems, totalling 5/100. The paper attributes this to BLIP-2's use of FLAN-T5-XXL as its language model, which is less capable than the Vicuna LLM used in MiniGPT-4, and notes that even after fine-tuning BLIP-2 on the authors' second-stage data it still generates short responses and fails to generalise to advanced tasks.