IC-428Qwen 1.5 appears to Pareto-dominate Pythia and LLaMA 2 on MMLU and GSM8K, but after adjusting for test task training all three model families exhibit equivalent scaling

Ricardo Dominguez-Olmedo, Florian E. Dorner, Moritz Hardt

SourceTraining on the Test Task Confounds Evaluation and Emergence

The paper compares three model families that likely train on the test task to very different extents: Pythia (trained on The Pile, unlikely to contain test task data), LLaMA 2 (trained mostly on web data), and Qwen 1.5 (explicitly includes instruction data in pretraining). Without adjustment, Qwen 1.5 Pareto-dominates both other families while Pythia models perform near random chance. After fine-tuning all models on the same task-relevant data, all three families show remarkably similar scaling trends, and none appears superior. This demonstrates that apparent family-level superiority is a confound of differential test task training.

Evidence
interventional
Caveat
The paper does not report specific accuracy numbers for each family in the text; the comparison is presented qualitatively in Figure 6. The conclusion depends on the assumption that the three families differ primarily in test task training exposure.
Model
Pythia, Llama 2 / Llama 2 base, Qwen1.5
Concepts
Method artefact
Datasets
MMLU / MMLU-Math [eval], GSM8K [eval]
Related findings
IC-427, IC-429
Extraction
automatic-extraction