IC-765BLIP-2, BLIP, InstructBLIP, Bard, and BingChat fall short of human accuracy in answering questions about Midjourney-generated images

Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D. Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, Yejin Choi

SourceThe Generative AI Paradox: “What It Can Create, It May Not Understand”

In the interrogative evaluation for vision, Midjourney generates images from text prompts, and understanding models are asked questions about the content of those generated images. All tested models (BLIP, BLIP-2 variants, InstructBLIP variants, Bard, BingChat) score below human accuracy across COCO, PaintSkill, DrawBench, and Parti subsets. Notably, the performance gap is smaller for simpler models (BLIP-2) than for advanced multimodal LLMs (Bard, BingChat), which have some visual understanding abilities but still struggle with simple questions about generated images.

Evidence
correlational
Key metric
blip2-flan-t5-xxl: 91.18% (COCO), 88.59% (PaintSkill), 85.31% (DrawBench), 90.56% (Parti) vs human 95.88%, 97.72%, 96.32%, 96.83%; bard: 74.02%, 66.28%, 56.33%, 59.42%; bingchat: 80.49%, 87.20%, 80.68%, 87.20%
Caveat
Bard and BingChat can occasionally refuse to answer when the image contains people; results for these models are on a subset where they provide a reasonable answer. All results are zero-shot. 1,871 image-question pairs were used in total.
Model
BLIP-2, BLIP, InstructBLIP, Bard, BingChat
Concepts
Failure mode
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval], DrawBench [eval], Parti [eval]
Methods
Zero-shot prompting [primary]
Related work
Li et al. 2023a (Your Diffusion Model is Secretly a Zero-Shot Classifier) [context]
Related findings
IC-762, IC-763, IC-764
Extraction
automatic-extraction