IC-633BLIP-2, LLaVA, and mPLUG-Owl show a trade-off between caption length and hallucination rate on COCO

Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny

SourceMiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

The paper measures the CHAIRi hallucination metric and average caption length for three released vision-language models on COCO. BLIP-2 produces very short captions (6.5 words average) with minimal hallucination (CHAIRi 1.3). LLaVA produces longer captions (90.7 words) with moderate hallucination (CHAIRi 18.8). mPLUG-Owl produces the longest captions (98.5 words) with the highest hallucination (CHAIRi 30.2). The paper notes that longer captions tend to have higher hallucination rates, and that BLIP-2's low hallucination comes at the cost of covering fewer objects.

Evidence
correlational
Key metric
BLIP-2: CHAIRi 1.3, avg length 6.5; mPLUG-Owl: CHAIRi 30.2, avg length 98.5; LLaVA: CHAIRi 18.8, avg length 90.7
Model
BLIP-2, LLaVA, mPLUG-Owl
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval]
Methods
CHAIR / CHAIRi / CHAIRS [eval]
Related findings
IC-632, IC-634, IC-635
Extraction
automatic-extraction