IC-1340Chinchilla 70B produces coherent autoregressive continuations of text, image, and audio data when used as a compressor, outperforming gzip in sample quality

Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness

SourceLanguage Modeling Is Compression

The paper demonstrates the prediction-compression equivalence by using compressors as generative models: the conditional probability of the next symbol is derived from the change in compressed length. When conditioning on the first half of each row of an ImageNet image (250 pixels) and sampling the remaining 250 pixels autoregressively, Chinchilla 70B produces visually appropriate continuations that degrade with length, while gzip produces much noisier completions. Similar qualitative comparisons are made for text (enwik9) and audio (Librispeech), where Chinchilla 70B samples are significantly more coherent than gzip's.

Evidence
observational
Caveat
The comparison is qualitative ('judged qualitatively'); no quantitative generation metric is reported. The paper notes that gzip samples are biased because its internal dictionary of tokens can be longer than one byte, and that looking multiple steps ahead would improve results. The image generation treats rows as independent, which is an oversimplification of natural image statistics.
Model
Chinchilla
Datasets
ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval], enwik9 [eval], LibriSpeech [eval]
Methods
Arithmetic Coding [primary], gzip [compared-to]
Related findings
IC-1338, IC-1339
Extraction
automatic-extraction