IC-489State-of-the-art MLLMs fail at multi-step visual analogical reasoning, with best accuracy at 13% (Llama 3.2) on VOILA-WD and 29% (GPT-4o) on VOILA-ND, far below human performance of 71% and 70%
Nilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang
Seven released MLLMs were evaluated on the VOILA benchmark, which requires models to infer a relational rule from a reference image pair and apply it to a new image pair to predict a fourth image. Performance degrades sharply across the four reasoning steps: models achieve 50-79% on image description, drop to 12-43% on relationship identification, and fall to 0.4-29% on applying the relationship. The gap between the best model and human participants is approximately 58 percentage points on VOILA-WD and 40 on VOILA-ND. The distraction rule in VOILA-WD further reduces accuracy for all models except Llama 3.2, which handles it better than other models.
Evaluation uses GPT-4o as an automated judge with up to 10% error rate per step; image generation step evaluated only on GPT-4o, Seed-LLaMA, and Emu2 due to limited image generation capabilities of other models.