IC-429The point of emergence for MMLU shifts from approximately 1.3×10²² flops to 5.6×10²⁰ flops as models train on 64,000 task-relevant examples, and the log-linear fit R² improves from 0.632 to 0.950
Ricardo Dominguez-Olmedo, Florian E. Dorner, Moritz Hardt
The paper evaluates older models (pre-November 2023) at intermediate checkpoints during fine-tuning on task-relevant data. As the number of training examples increases from 0 to 64,000, the compute threshold at which emergence occurs (ce) drops from 1.3×10²² flops (scale of Pythia 6.9B) to 5.6×10²⁰ flops (scale of Pythia 410M). Simultaneously, the R² of the log-linear fit between pretraining compute and benchmark accuracy improves from 0.632 to 0.950. Similar results are observed for GSM8K, where ce drops from 2.2×10²² to 1.3×10²¹ and R² improves from 0.515 to 0.857. This shows that emergence is not an intrinsic property of the model but is modulated by the amount of test task training.
Results are for models trained before November 2023 only. The emergence threshold is estimated from a regression fit and depends on the chosen functional form. The authors note that Brier score does not eliminate emergence for MMLU, so the effect is not purely a metric artefact.