Theoretical analysis demonstrates algorithmic information limits in large language models, indicating that generative systems cannot exceed the empirical information of their training data.
This is a position paper. It advances no new mathematical results; its purpose is to establish a constraint that subsequent theoretical work on generative models can take as a premise. The claim is that a large language model (LLM), understood as a computable function —deterministic or stochastic— applied to a fixed corpus of human data, cannot generate net algorithmic information exceeding what that corpus contains plus the complexity of the algorithm itself. Drawing on three independent frameworks —Hartley’s combinatorial measure, Shannon’s statistical theory, and the algorithmic information theory of Kolmogorov, Chaitin and Solomonoff— and resting centrally on Levin’s Information Conservation Laws (1974) and the Data Processing Inequality (DPI), we argue that any system AI = LLM(H) is bounded above by K(H) +K(LLM) +O(1), and that its mutual information with physical reality R cannot exceed that already present in H. Model collapse is derived as a corollary: monotone degradation under iteration of lossy channels. We state the condition under which an artificial system generates genuinely new empirical information —the acquisition of a measurement channel to reality external to the training corpus— and close by setting out what follows for subsequent theoretical work if the constraint is accepted.
No takes yet. Share an insight, caveat, or question.
Juan Pablo Venegas Padilla (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: