PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 19, 202328 citationsOpen Access

Language Modeling Is Compression

GDGrégoire DelétangARAnian RuossPDPaul-Ambroise Duquenne

Key Points

  • The aim is to evaluate the compression capabilities of large language models and explore insights derived from their predictive functions.
  • Assessment of large language models' compression capabilities across various data types.
  • Comparison of compression efficiency against domain-specific compressors.
  • Evaluation of implications on scaling laws, tokenization, and in-context learning.
  • Chinchilla 70B compresses ImageNet patches to 43.4% of their raw size.
  • LibriSpeech samples are compressed to 16.4% of their original size.
  • Large language models outperform traditional compressors like PNG and FLAC.

Abstract

It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Delétang et al. (2023) studied this question.

synapsesocial.com/papers/69b03ca398a0803b6cb32c12https://doi.org/10.48550/arxiv.2309.10668
Ask AI
Helpful
Bookmark
Share
View Full Paper