Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MathVista
anchor
Findings
IC-276
All 14 evaluated VLMs show a large gap between average-case and worst-case accuracy on DynaMath variants, with worst-case at or below 50% of average-case, and the failures are systematic rather than random
[compared-to]
IC-440
GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet show up to 25% skill-level accuracy gaps despite overall accuracies within 0.4% of each other
[eval]
IC-441
Skill-level improvements between model releases are highly uneven, with Claude 3.5 Sonnet gaining ~50% over Claude 3 Opus on law skills while Gemini improved most in math and science
[eval]
IC-442
Routing each evaluation instance to the model strongest on its relevant skills yields a 3.2% accuracy gain over the best single model, with 3.5-6.8% gains on MMLU Pro
[eval]
IC-443
Model inconsistency on probing questions negatively correlates with skill-slice accuracy (r = -0.675), with models contradicting themselves more often on skills where they perform poorly
[eval]