The paper evaluates three frontier models on 46k instances across 12 benchmarks, grouping instances into skill-slices (sets sharing a relevant skill). While aggregate accuracy is nearly identical, per-slice accuracy reveals stark trade-offs: Gemini 1.5 Pro is 18% more accurate at computing molar mass but 19% less accurate at applying constitutional law; GPT-4o is 16.7% more accurate at traffic signal identification. GPT-4o's relative strengths are visual skills, Gemini 1.5 Pro's are math and science, and Claude 3.5 Sonnet's are legal reasoning.
Evidence
correlational
Key metric
overall accuracies within 0.4%; differences as high as 25% for certain skills; gemini 1.5 pro 18% more accurate in computing molar mass, 19% less accurate in applying constitutional law; gpt-4o 16.7% more accurate in traffic signal identification
Caveat
The authors note they cannot offer explanations for the strengths and weaknesses observed, as they have limited visibility into the training data and processes of these private models.