PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 27, 2026Publications0 citationsOpen Access

Large Language Models in Peer Review: Decision Alignment, Review-Text Characteristics, and Human–AI Aggregation at ICLR 2025

View Full Paper
ZYZhihe YangXZXiaoyu ZhouHWHongsa Wang

Key Points

  • To assess the viability of large language models as autonomous peer reviewers by comparing their scoring behavior, review characteristics, and decision alignment against human reviewers at ICLR 2025.
  • Analyzed 2,401 human reviews alongside 7,203 context-isolated API reviews generated by Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview.
  • Used a 0.1-point grid search with repeated stratified cross-validation to calibrate decision thresholds, validating them on a balanced sample of 300 ICLR 2024 papers.
  • Evaluated text characteristics, eight-gram containment for data contamination, and independent human coding agreement of research type and primary field (Cohen’s kappa = 0.774 and 0.714).
  • Raw LLM evaluations demonstrated systematic score compression and leniency, yielding calibrated acceptance thresholds of 6.2 (Claude), 6.3 (GPT), and 6.7 (Gemini).
  • Calibrated decision-agreement accuracies reached 0.927 for Claude, 0.913 for GPT, and 0.930 for Gemini on the ICLR 2024 validation dataset.
  • Human-containing aggregation rules achieved higher agreement with final conference decisions than AI-only rules, while LLM reviews showed uneven critical-section lengths despite structural completeness.

Abstract

Large language models (LLMs) are increasingly employed in scholarly peer review, yet their suitability as autonomous evaluators remains uncertain. Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation. Raw LLM scores showed systematic leniency and score compression. A 0.1-point grid search identified thresholds of 6.2, 6.3, and 6.7 for Claude, GPT, and Gemini, respectively; repeated stratified cross-validation reproduced these thresholds. When applied without retuning to a stratified balanced sample of 300 ICLR 2024 papers, decision-agreement accuracy was 0.927, 0.913, and 0.930. Independent human coding of research type and primary field showed substantial pre-adjudication agreement (Cohen’s kappa = 0.774 and 0.714), and the recalculated analyses did not support H3. Review-text indicators showed similar structural completeness across sources but uneven critical-section length; these descriptive measures do not establish review quality. Human-containing aggregation rules showed higher agreement with conference decisions than corresponding AI-only rules, without establishing independent review quality or causal complementarity. A textual-overlap check found very low exact eight-gram containment, and manual inspection of the highest-similarity 1% found shared manuscript content or domain terminology rather than reviewer-specific evaluative language; possible prior exposure nevertheless could not be excluded.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2026) studied this question.

synapsesocial.com/papers/6a8fea0110c91c1e92621eechttps://doi.org/10.3390/publications14030055
Ask AI
Helpful
Bookmark
Share
View Full Paper