SourceMiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
The paper measures the CHAIRi hallucination metric and average caption length for three released vision-language models on COCO. BLIP-2 produces very short captions (6.5 words average) with minimal hallucination (CHAIRi 1.3). LLaVA produces longer captions (90.7 words) with moderate hallucination (CHAIRi 18.8). mPLUG-Owl produces the longest captions (98.5 words) with the highest hallucination (CHAIRi 30.2). The paper notes that longer captions tend to have higher hallucination rates, and that BLIP-2's low hallucination comes at the cost of covering fewer objects.