Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Training on the Test Task Confounds Evaluation and Emergence
2025-01-22
· ICLR 2025 Oral ·
anchor
Findings
IC-427
Newer base models (post-November 2023) outperform older ones by 7.3 points on MMLU and 19.1 points on GSM8K controlling for pretraining compute, but this gap vanishes after fine-tuning all models on the same task-relevant data
IC-428
Qwen 1.5 appears to Pareto-dominate Pythia and LLaMA 2 on MMLU and GSM8K, but after adjusting for test task training all three model families exhibit equivalent scaling
IC-429
The point of emergence for MMLU shifts from approximately 1.3×10²² flops to 5.6×10²⁰ flops as models train on 64,000 task-relevant examples, and the log-linear fit R² improves from 0.632 to 0.950