PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 20251 citationsOpen Access

A suite of LMs comprehend puzzle statements as well as humans

View Full Paper
AGAdele Ε. GoldbergSRSupantho RakshitJHJennifer Hu

Key Points

  • Human performance in comprehension tasks declines when rereading is restricted, dropping to 73%.
  • LLM models like Falcon-180B-Chat and GPT-4 demonstrate 76% and 81% accuracy respectively, outperforming humans.
  • Analyses affirm that both humans and models face challenges with queries involving reciprocal actions, indicating shared sensitivities.
  • Findings argue for refined experimental designs in evaluating LLMs, questioning assumptions about their comprehension capabilities.

Abstract

Recent claims suggest that large language models (LMs) underperform humans in comprehending minimally complex English statements (Dentella et al., 2024). Here, we revisit those findings and argue that human performance was overestimated, while LLM abilities were underestimated. Using the same stimuli, we report a preregistered study comparing human responses in two conditions: one allowed rereading (replicating the original study), and one that restricted rereading (a more naturalistic comprehension test). Human accuracy dropped significantly when rereading was restricted (73%), falling below that of Falcon-180B-Chat (76%) and GPT-4 (81%). The newer GPT-o1 model achieves perfect accuracy. Results further show that both humans and models are disproportionately challenged by queries involving potentially reciprocal actions (e.g., kissing), suggesting shared pragmatic sensitivities rather than model-specific deficits. Additional analyses using Llama-2-70B log probabilities, a recoding of open-ended model responses, and grammaticality ratings of other sentences reveal systematic underestimation of model performance. We find that GPT-4o can align with either naive or expert grammaticality judgments, depending on prompt framing. These findings underscore the need for more careful experimental design and coding practices in LLM evaluation, and they challenge the assumption that current models are inherently weaker than humans at language comprehension.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Goldberg et al. (2025) studied this question.

synapsesocial.com/papers/68e82b12e7fc21a3005002b2https://doi.org/10.48550/arxiv.2505.08996
Ask AI
Helpful
Bookmark
Share
View Full Paper