PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20251 citationsOpen Access

DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process

View Full Paper
MZMin ZhuYWYixuan WengLYLinyi Yang

Key Points

  • DeepReviewer-14B outperformed CycleReviewer-70B with fewer tokens, achieving win rates of 88.21% against GPT-o1.
  • In evaluations, DeepReviewer-14B surpassed competitors with a win rate of 80.20% against DeepSeek-R1, showcasing its effectiveness.
  • DeepReview employs a multi-stage framework that combines literature retrieval and evidence-based argumentation for improved analysis.
  • This work establishes a new benchmark for LLM-based paper review, making the code and dataset publicly available.

Abstract

Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review. However, existing LLM-based review systems face significant challenges, including limited domain expertise, hallucinated reasoning, and a lack of structured evaluation. To address these limitations, we introduce DeepReview, a multi-stage framework designed to emulate expert reviewers by incorporating structured analysis, literature retrieval, and evidence-based argumentation. Using DeepReview-13K, a curated dataset with structured annotations, we train DeepReviewer-14B, which outperforms CycleReviewer-70B with fewer tokens. In its best mode, DeepReviewer-14B achieves win rates of 88.21\% and 80.20\% against GPT-o1 and DeepSeek-R1 in evaluations. Our work sets a new benchmark for LLM-based paper review, with all resources publicly available. The code, model, dataset and demo have be released in http://ai-researcher.net.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2025) studied this question.

synapsesocial.com/papers/68da58c9c1728099cfd10b0chttps://doi.org/10.48550/arxiv.2503.08569
Ask AI
Helpful
Bookmark
Share
View Full Paper