IC-334GPT-4V produces more poetic, emotion-focused image captions compared to Gemini-1.5-Flash's literal descriptions, with 99% model-matching accuracy on COCO

Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez

SourceVibeCheck: Discover and Quantify Qualitative Differences in Large Language Models

On 1000 COCO images, VibeCheck compares captions generated by GPT-4V and Gemini-1.5-Flash. GPT-4V uses more poetic language, structures captions as dynamic stories, infers personality and emotions of subjects, and emphasizes mood and atmosphere (sep 0.63). Gemini-1.5-Flash sticks to straightforward, literal descriptions focusing on identifying objects. The top 10 vibes achieve 99.13% model-matching accuracy and 89.02% preference-prediction accuracy. GPT-4V has an 80% win rate in LLM-judged preference.

Evidence
correlational
Key metric
Model matching accuracy 99.13%; preference prediction accuracy 89.02%; GPT-4V win rate 80%; separability scores: color and atmosphere description 0.63, emotion and relationships 0.60, background details 0.56, sense of space 0.42
Caveat
Captions were compared without the image due to cost. The paper notes the VibeCheck framework can be adapted to the multimodal setting but this was not done here.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Flash
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval], ShareGPT4V [source]
Related work
ShareGPT4V [context]
Related findings
IC-332, IC-333
Extraction
automatic-extraction