IC-869Models ranking highly on popular LLM leaderboards perform worse than Llama-2-70b-chat on Skill-Mix, suggesting cramming for the leaderboard at the expense of general-purpose text generation

Dingli Yu, Simran Kaur, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora

SourceSKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models

Falcon-180b-chat, Xwin-LM-70b-v0.1, Qwen-14b-chat, and Mistral-7b-instruct-v0.1 all rank above Llama-2-70b-chat on the Open LLM Leaderboard or AlpacaEval, yet all perform worse than Llama-2-70b-chat on Skill-Mix under both GPT-4 and Llama-2 grading. TigerBot-70b-chat, also ranked above Llama-2-70b-chat on the Open LLM Leaderboard, performs even worse than Llama-2-13b-chat on Skill-Mix. The authors interpret this as evidence that training on data similar to leaderboard benchmarks ('cramming') significantly harmed general-purpose text skills.

Evidence
correlational
Key metric
Ratio of full marks (GPT-4 grading, k=2): Falcon-180b-chat .27, TigerBot-70b-chat .06, Qwen-14b-chat .16, Mistral-7b-instruct-v0.1 .05, versus Llama-2-70b-chat .28
Caveat
The paper does not directly measure whether these models were trained on leaderboard data; the 'cramming' interpretation is inferential. Falcon-180b-chat was evaluated on only 30 combinations.
Model
Falcon Falcon-180B-Chat, Xwin-LM-70B-v0.1, Qwen Qwen-14B-Chat, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.1, TigerBot-70B-Chat, Llama 2 / Llama 2 base
Datasets
Open LLM Leaderboard [eval], AlpacaEval [eval]
Related work
Open LLM Leaderboard [compared-to], AlpacaEval [compared-to]
Related findings
IC-868, IC-870, IC-871
Extraction
automatic-extraction