IC-869Models ranking highly on popular LLM leaderboards perform worse than Llama-2-70b-chat on Skill-Mix, suggesting cramming for the leaderboard at the expense of general-purpose text generation
Falcon-180b-chat, Xwin-LM-70b-v0.1, Qwen-14b-chat, and Mistral-7b-instruct-v0.1 all rank above Llama-2-70b-chat on the Open LLM Leaderboard or AlpacaEval, yet all perform worse than Llama-2-70b-chat on Skill-Mix under both GPT-4 and Llama-2 grading. TigerBot-70b-chat, also ranked above Llama-2-70b-chat on the Open LLM Leaderboard, performs even worse than Llama-2-13b-chat on Skill-Mix. The authors interpret this as evidence that training on data similar to leaderboard benchmarks ('cramming') significantly harmed general-purpose text skills.
Evidence
correlational
Key metric
Ratio of full marks (GPT-4 grading, k=2): Falcon-180b-chat .27, TigerBot-70b-chat .06, Qwen-14b-chat .16, Mistral-7b-instruct-v0.1 .05, versus Llama-2-70b-chat .28
Caveat
The paper does not directly measure whether these models were trained on leaderboard data; the 'cramming' interpretation is inferential. Falcon-180b-chat was evaluated on only 30 combinations.