Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Least-to-Most Prompting
anchor
Findings
IC-489
State-of-the-art MLLMs fail at multi-step visual analogical reasoning, with best accuracy at 13% (Llama 3.2) on VOILA-WD and 29% (GPT-4o) on VOILA-ND, far below human performance of 71% and 70%
[primary]
IC-490
GPT-4o can identify visual relationships at 97% accuracy when given ground-truth descriptions but drops to 17% when asked to apply known relationships to new visuals, revealing a specific bottleneck in relational transfer
[primary]
IC-491
Presenting three images as a single collage rather than sequentially reduces MLLM accuracy by approximately 40% on the relationship application step
[primary]
IC-492
Least-to-most prompting consistently improves MLLM accuracy on the relationship application step compared to direct answering, with GPT-4o improving from 0.9% to 6.44% on VOILA-WD
[builds-on]
IC-492
Least-to-most prompting consistently improves MLLM accuracy on the relationship application step compared to direct answering, with GPT-4o improving from 0.9% to 6.44% on VOILA-WD
[primary]