IC-441Skill-level improvements between model releases are highly uneven, with Claude 3.5 Sonnet gaining ~50% over Claude 3 Opus on law skills while Gemini improved most in math and science
Mazda Moayeri, Vidhisha Balachandran, Varun Chandrasekaran, Safoora Yousefi, Thomas FEL, Soheil Feizi, Besmira Nushi, Neel Joshi, Vibhav Vineet
The paper compares the last two releases in each of the GPT, Gemini, and Claude families on skill-slices. For all families, some skills show improvement more than twice the average while others show no improvement at all. The GPT and Claude families saw their largest gains in legal skills, with Claude 3.5 Sonnet improving over Claude 3 Opus by approximately 50% on numerous law skill-slices. Gemini 1.5 Pro improved most in math and science skills relative to Gemini 1.0 Pro.
Evidence
correlational
Key metric
claude 3.5 sonnet improving over claude opus 3.0 by a staggering ~50% along numerous law skill-slices
Caveat
Gemini 1.0 Pro is not multimodal, so results for Gemini are presented on 12k language-only evaluation instances rather than the full multimodal corpus.