PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 15, 2025Proceedings of the National Academy of Sciences37 citationsOpen Access

The simulation of judgment in LLMs

View Full Paper
ELEdoardo LoruJNJacopo NudoNMNiccolò Di Marco

Key Points

  • LLMs show consistent differences in evaluation strategies compared to human judgments, affecting how they interpret criteria.
  • Benchmarking utilized expert ratings from NewsGuard and Media Bias/Fact Check, showing different influences on evaluations.
  • The structured evaluation framework allowed direct comparisons, revealing the divergence in reasoning between LLMs and humans.
  • Findings suggest a shift toward pattern-based reasoning, prompting concerns regarding the reliability of LLMs in evaluative tasks.

Abstract

Large Language Models (LLMs) are increasingly embedded in evaluative processes, from information filtering to assessing and addressing knowledge gaps through explanation and credibility judgments. This raises the need to examine how such evaluations are built, what assumptions they rely on, and how their strategies diverge from those of humans. We benchmark six LLMs against expert ratings—NewsGuard and Media Bias/Fact Check—and against human judgments collected through a controlled experiment. We use news domains purely as a controlled benchmark for evaluative tasks, focusing on the underlying mechanisms rather than on news classification per se. To enable direct comparison, we implement a structured agentic framework in which both models and nonexpert participants follow the same evaluation procedure: selecting criteria, retrieving content, and producing justifications. Despite output alignment, our findings show consistent differences in the observable criteria guiding model evaluations, suggesting that lexical associations and statistical priors could influence evaluations in ways that differ from contextual reasoning. This reliance is associated with systematic effects: political asymmetries and a tendency to confuse linguistic form with epistemic reliability—a dynamic we term epistemia, the illusion of knowledge that emerges when surface plausibility replaces verification. Indeed, delegating judgment to such systems may affect the heuristics underlying evaluative processes, suggesting a shift from normative reasoning toward pattern-based approximation and raising open questions about the role of LLMs in evaluative processes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Loru et al. (2025) studied this question.

synapsesocial.com/papers/68eff7392ae617e5891a94f2https://doi.org/10.1073/pnas.2518443122
Ask AI
Helpful
Bookmark
Share
View Full Paper