Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-868
Skill-Mix performance degrades with increasing k, and within the Llama-2 family the saturation point increases with model size
IC-869
Models ranking highly on popular LLM leaderboards perform worse than Llama-2-70b-chat on Skill-Mix, suggesting cramming for the leaderboard at the expense of general-purpose text generation
IC-870
GPT-4's performance on Skill-Mix(k=5) and Skill-Mix(k=6) provides probabilistic evidence of generating novel skill-topic combinations not present in training data
IC-871
Llama-2-70b-chat as a grader is more generous than GPT-4 and systematically gives higher scores to Llama-2 family outputs