Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MMMU
anchor
Findings
IC-344
GPT-4o achieves the highest average score (64.62) on general document benchmarks, outperforming Qwen2-VL-72B (58.40) and GeminiPro-1.5 (57.05)
[eval]
IC-440
GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet show up to 25% skill-level accuracy gaps despite overall accuracies within 0.4% of each other
[eval]
IC-441
Skill-level improvements between model releases are highly uneven, with Claude 3.5 Sonnet gaining ~50% over Claude 3 Opus on law skills while Gemini improved most in math and science
[eval]
IC-442
Routing each evaluation instance to the model strongest on its relevant skills yields a 3.2% accuracy gain over the best single model, with 3.5-6.8% gains on MMLU Pro
[eval]
IC-443
Model inconsistency on probing questions negatively correlates with skill-slice accuracy (r = -0.675), with models contradicting themselves more often on skills where they perform poorly
[eval]