In the interrogative evaluation for vision, Midjourney generates images from text prompts, and understanding models are asked questions about the content of those generated images. All tested models (BLIP, BLIP-2 variants, InstructBLIP variants, Bard, BingChat) score below human accuracy across COCO, PaintSkill, DrawBench, and Parti subsets. Notably, the performance gap is smaller for simpler models (BLIP-2) than for advanced multimodal LLMs (Bard, BingChat), which have some visual understanding abilities but still struggle with simple questions about generated images.
Bard and BingChat can occasionally refuse to answer when the image contains people; results for these models are on a subset where they provide a reasonable answer. All results are zero-shot. 1,871 image-question pairs were used in total.