IC-186The memorization status of sequences in Pythia-1B and Amber-7B is stationary throughout training: KL-LD fluctuations are mean-reverting with fixed variance, rejecting a random-walk model with p < 10⁻⁸
Sunny Duan, Mikail Khona, Abhiram Iyer, Rylan Schaeffer, Ila R Fiete
The paper tracks KL-LD for individual sequences (z-complexity > 0.8, encountered only once) across checkpoints from step 10k to 43k (Pythia-1B) and revisions 100 to 350 (Amber-7B). The distribution of KL-LD at each checkpoint is the same across all checkpoints, the variance does not grow over time, and a variance ratio test rejects the random-walk null hypothesis with p < 10⁻⁸ for both models. Changes between consecutive checkpoints are symmetric and roughly Laplace-distributed, indicating sequences become memorized as often as they are forgotten.
Evidence
correlational
Key metric
variance ratio test rejects random walk null hypothesis with p < 10⁻⁸ for both Pythia-1B and Amber-7B
Caveat
Analysis restricted to sequences with z-complexity > 0.8 and no sub-sequence match of length 30 or longer; only a subset of the full training trajectory is examined (steps 10k–43k for Pythia-1B).