IC-185In Pythia-1B and Amber-7B, the probability of memorizing a training sequence scales log-linearly with both the number of repetitions in the corpus and the z-complexity of the sequence

Sunny Duan, Mikail Khona, Abhiram Iyer, Rylan Schaeffer, Ila R Fiete

SourceUncovering Latent Memories in Large Language Models

The paper measures KL-Levenshtein distance (KL-LD) between the model's greedy continuation and the true continuation for sequences of varying repetition count and z-complexity (ratio of compressed to original length). For both Pythia-1B and Amber-7B, sequences with more repetitions and lower z-complexity are more easily memorized, and the relationship is log-linear in both factors. Lower-complexity strings require fewer repeats to reach the same memorization level.

Evidence
correlational
Caveat
The paper notes this extends prior work (Carlini et al. 2020) and the specific quantitative curves are shown in figures rather than tabulated in the text.
Model
Pythia 1B, Amber-7B
Datasets
The Pile [source]
Related findings
IC-186, IC-187
Extraction
automatic-extraction