Randomized trial demonstrates improved literature survey generation performance in scientific publications, suggesting a scalable solution for information overload.
The escalating volume and complexity of scientific publications have intensified the demand for intelligent literature survey tools. Current systems, whether based on large language models (LLMs) or retrieval augmented generation (RAG), often struggle with terminological accuracy and semantic coherence, limiting their practical adoption. To address this, we propose a novel LLM-driven framework that synergizes multi-agent collaboration with RAG to enhance both the structural coherence of full-document surveys and the factual reliability of generated content. On the SciReviewGen benchmark, our model achieves Recall-Oriented Understudy for Gisting Evaluation (ROUGE)-1/2/L scores of 42.84/16.57/18.25, outperforming strong baselines including Query-weighted Fusion-in-Decoder (QFiD), Fusion-in-Decoder (FiD), BigBird, GPT-4o-mini, and Llama4-17B-16E. The framework also excels in human-aligned evaluation, attaining an LLM-as-judge score of 0.70, surpassing QFiD by +0.29. Ablation studies confirm the critical role of both RAG and multi-agent design: removing RAG reduces ROUGE-2/L by 2.28/3.19, while single-agent ablation further decreases ROUGE-1 by 1.15. Our work not only advances the state of the art in automated literature synthesis but also offers a scalable solution to mounting scholarly information overload.
No takes yet. Share an insight, caveat, or question.
Qi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: