Language Modeling Is Compression
First page

Paper summary
DeepMind empirically revisits the theoretical equivalence between prediction and compression, applied to modern LLMs.
Ask this paper
01
Theoretical equivalence: Reminds that optimal compression and optimal prediction are duals - a good language model is implicitly a powerful compressor.
02
ImageNet compression: Chinchilla 70B compresses ImageNet patches to 43.4% of raw size, better than domain-specific codecs like PNG.
03
LibriSpeech compression: Compresses LibriSpeech samples to 16.4% of raw size, beating FLAC and gzip on audio data despite never being trained on audio.
04
Cross-modal generalization: Shows LLMs work as general-purpose compressors across text, image, and audio - a striking demonstration of in-context learning's reach.