PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 9, 20251 citationsOpen Access

Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge

View Full Paper
DPDrago PlečkoPOPatrik OkanovicTHTorsten Hoefler

Key Points

  • Language models demonstrate limited ability to learn and internalize real-world statistics, struggling particularly with empirical distributions.
  • Results indicate poor performance across multiple domains, including education and social behavior, challenging their universal application claims.
  • Developed a benchmark to assess language models' knowledge of observational and probabilistic distributions within real-world contexts.
  • Findings highlight the implications for the understanding of artificial intelligence in relation to causal hierarchies and distributional learning limits.

Abstract

Artificial intelligence (AI) systems hold great promise for advancing various scientific disciplines, and are increasingly used in real-world applications. Despite their remarkable progress, further capabilities are expected in order to achieve more general types of intelligence. A critical distinction in this context is between factual knowledge, which can be evaluated against true or false answers (e.g., "what is the capital of England?"), and probabilistic knowledge, reflecting probabilistic properties of the real world (e.g., "what is the sex of a computer science graduate in the US?"). In this paper, our goal is to build a benchmark for understanding the capabilities of LLMs in terms of knowledge of probability distributions describing the real world. Given that LLMs are trained on vast amounts of text, it may be plausible that they internalize aspects of these distributions. Indeed, LLMs are touted as powerful universal approximators of real-world distributions. At the same time, classical results in statistics, known as curse of dimensionality, highlight fundamental challenges in learning distributions in high dimensions, challenging the notion of universal distributional learning. In this work, we develop the first benchmark to directly test this hypothesis, evaluating whether LLMs have access to empirical distributions describing real-world populations across domains such as economics, health, education, and social behavior. Our results demonstrate that LLMs perform poorly overall, and do not seem to internalize real-world statistics naturally. When interpreted in the context of Pearl's Causal Hierarchy (PCH), our benchmark demonstrates that language models do not contain knowledge on observational distributions (Layer 1 of PCH), and thus the Causal Hierarchy Theorem implies that interventional (Layer 2) and counterfactual (Layer 3) knowledge of these models is also limited.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Plečko et al. (2025) studied this question.

synapsesocial.com/papers/690fdcdaf60c54d04ea38057https://doi.org/10.48550/arxiv.2511.03070
Ask AI
Helpful
Bookmark
Share
View Full Paper