Large language models (LLMs) are increasingly employed in scholarly peer review, yet their suitability as autonomous evaluators remains uncertain. Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation. Raw LLM scores showed systematic leniency and score compression. A 0.1-point grid search identified thresholds of 6.2, 6.3, and 6.7 for Claude, GPT, and Gemini, respectively; repeated stratified cross-validation reproduced these thresholds. When applied without retuning to a stratified balanced sample of 300 ICLR 2024 papers, decision-agreement accuracy was 0.927, 0.913, and 0.930. Independent human coding of research type and primary field showed substantial pre-adjudication agreement (Cohen’s kappa = 0.774 and 0.714), and the recalculated analyses did not support H3. Review-text indicators showed similar structural completeness across sources but uneven critical-section length; these descriptive measures do not establish review quality. Human-containing aggregation rules showed higher agreement with conference decisions than corresponding AI-only rules, without establishing independent review quality or causal complementarity. A textual-overlap check found very low exact eight-gram containment, and manual inspection of the highest-similarity 1% found shared manuscript content or domain terminology rather than reviewer-specific evaluative language; possible prior exposure nevertheless could not be excluded.
Yang et al. (Tue,) studied this question.