IC-440GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet show up to 25% skill-level accuracy gaps despite overall accuracies within 0.4% of each other

Mazda Moayeri, Vidhisha Balachandran, Varun Chandrasekaran, Safoora Yousefi, Thomas FEL, Soheil Feizi, Besmira Nushi, Neel Joshi, Vibhav Vineet

SourceUnearthing Skill-level Insights for Understanding Trade-offs of Foundation Models

The paper evaluates three frontier models on 46k instances across 12 benchmarks, grouping instances into skill-slices (sets sharing a relevant skill). While aggregate accuracy is nearly identical, per-slice accuracy reveals stark trade-offs: Gemini 1.5 Pro is 18% more accurate at computing molar mass but 19% less accurate at applying constitutional law; GPT-4o is 16.7% more accurate at traffic signal identification. GPT-4o's relative strengths are visual skills, Gemini 1.5 Pro's are math and science, and Claude 3.5 Sonnet's are legal reasoning.

Evidence
correlational
Key metric
overall accuracies within 0.4%; differences as high as 25% for certain skills; gemini 1.5 pro 18% more accurate in computing molar mass, 19% less accurate in applying constitutional law; gpt-4o 16.7% more accurate in traffic signal identification
Caveat
The authors note they cannot offer explanations for the strengths and weaknesses observed, as they have limited visibility into the training data and processes of these private models.
Model
GPT-4o, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Claude 3.5 Sonnet
Datasets
MMLU Pro [eval], MMMU [eval], MathVista [eval], MMC [eval], MMVP [eval], MMBench [eval], MMT-Bench [eval], MME [eval], MM-VET [eval], Seed-Bench / SeedBench [eval], Vibe-Eval [eval]
Related findings
IC-441, IC-442, IC-443
Extraction
automatic-extraction