IC-1338Chinchilla 70B and Llama 2 7B, trained primarily on text, compress ImageNet patches and Librispeech audio better than domain-specific compressors PNG and FLAC
Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness
The paper uses arithmetic coding to turn pretrained language models into lossless compressors and evaluates them on 1 GB datasets of text, image, and audio. Chinchilla 70B, which was trained on a mix of internet text and books, achieves raw compression rates of 48.0% on ImageNet patches and 21.0% on Librispeech audio, outperforming the domain-specific PNG compressor (58.5% on ImageNet) and FLAC compressor (30.3% on Librispeech). Llama 2 7B similarly achieves 53.4% on ImageNet and 23.1% on Librispeech. These models have not been explicitly trained on image or audio data, so their cross-modal compression is attributed to in-context learning.
The adjusted compression rate (accounting for model parameter size in float16) is extremely poor for these models (e.g., 14048.0% for Chinchilla 70B on ImageNet), meaning the raw compression advantage only holds when model size is not included in the code length. The paper notes it is possible, though unlikely, that some image or audio data was encoded into text in the training corpus.