IC-1339Chinchilla 1B's compression rate improves with increasing sequence length across text, image, and audio, demonstrating in-context learning without gradient updates

Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness

SourceLanguage Modeling Is Compression

The paper measures the compression rate of Chinchilla 1B, gzip, and a small transformer over subsequences of increasing length (up to 2048 bytes), averaged over 100 sequences per dataset. For all three data modalities (enwik9, ImageNet, Librispeech), the compression rate decreases (improves) as the sequence length grows, indicating the model is learning data statistics in-context. Chinchilla 1B achieves the best compression rates across all modalities and sequence lengths compared to gzip and the small transformer.

Evidence
correlational
Caveat
The results are from a plot (Figure 4) with no tabulated values; the improvement is described qualitatively as 'decrease quickly with increasing sequence length.' The small transformer used for comparison was trained by the authors on enwik8.
Model
Chinchilla
Datasets
enwik9 [eval], ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval], LibriSpeech [eval]
Methods
Arithmetic Coding [primary], gzip [compared-to]
Related findings
IC-1338, IC-1340
Extraction
automatic-extraction