IC-868Skill-Mix performance degrades with increasing k, and within the Llama-2 family the saturation point increases with model size

Dingli Yu, Simran Kaur, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora

SourceSKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models

The paper evaluates ten instruction-tuned models on Skill-Mix(k) for k=2 through k=6, where the model must generate text illustrating k randomly selected skills in a given topic. Performance drops sharply as k increases: most models saturate by k=3 under GPT-4 grading. Within the Llama-2 family, which shares training data and methodology, the saturation point rises with size: Llama-2-7b-chat saturates at k=2, while Llama-2-13b-chat and Llama-2-70b-chat saturate at k=3. GPT-4 is the only model that maintains reasonable performance through k=5 and k=6.

Evidence
correlational
Key metric
Ratio of full marks (GPT-4 grading): Llama-2-7b-chat k=2 .05, k=3 .00; Llama-2-13b-chat k=2 .22, k=3 .05; Llama-2-70b-chat k=2 .28, k=3 .02; GPT-4 k=2 .94, k=3 .68, k=4 .52, k=5 .36, k=6 .29
Caveat
Falcon-180b-chat was evaluated on only 30 combinations with GPT-4 grading due to computation limits, compared to 100 for other models.
Model
Llama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Falcon Falcon-180B-Chat, Xwin-LM-70B-v0.1, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.1, Qwen Qwen-14B-Chat, TigerBot-70B-Chat
Concepts
Scale-dependent behaviour
Related work
Arora & Goyal (2023) [builds-on]
Related findings
IC-869, IC-870, IC-871
Extraction
automatic-extraction