IC-490GPT-4o can identify visual relationships at 97% accuracy when given ground-truth descriptions but drops to 17% when asked to apply known relationships to new visuals, revealing a specific bottleneck in relational transfer

Nilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang

SourceVOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning

An ablation study on GPT-4o using VOILA-WD with L2M prompting provided ground-truth image descriptions at phase 2, after which GPT-4o identified relationships at 97% accuracy. However, when ground-truth relationships were provided at phase 3, GPT-4o's accuracy on applying those relationships to the third image dropped to 17%, well below the 71% human baseline. A complementary ablation on VOILA-ND showed GPT-4o achieves 22% accuracy with three sequential images but 49% with text descriptions of the same images, indicating a modality gap in visual relational reasoning.

Evidence
correlational
Key metric
GPT-4o identifying relations with GT descriptions: 97%; GPT-4o applying relations with GT relationships: 17%; GPT-4o with images (VOILA-ND): 22%; GPT-4o with text (VOILA-ND): 49%
Caveat
Ablation conducted only on GPT-4o; the 17% figure is on VOILA-WD which includes distraction rules, making the task harder than VOILA-ND.
Model
GPT-4o
Concepts
Failure mode
Methods
Least-to-Most Prompting [primary]
Related findings
IC-489, IC-491, IC-492
Extraction
automatic-extraction