IC-185In Pythia-1B and Amber-7B, the probability of memorizing a training sequence scales log-linearly with both the number of repetitions in the corpus and the z-complexity of the sequence
Sunny Duan, Mikail Khona, Abhiram Iyer, Rylan Schaeffer, Ila R Fiete
The paper measures KL-Levenshtein distance (KL-LD) between the model's greedy continuation and the true continuation for sequences of varying repetition count and z-complexity (ratio of compressed to original length). For both Pythia-1B and Amber-7B, sequences with more repetitions and lower z-complexity are more easily memorized, and the relationship is log-linear in both factors. Lower-complexity strings require fewer repeats to reach the same memorization level.
Evidence
correlational
Caveat
The paper notes this extends prior work (Carlini et al. 2020) and the specific quantitative curves are shown in figures rather than tabulated in the text.