Benchmark evaluation reveals improved accuracy across multi-hop reasoning tasks using cooperative multi-agent architectures, highlighting the utility of modular retrieval-augmented generation.
Knowledge-intensive Natural Language Processing (NLP) tasks, such as open-domain question answering, multi-hop reasoning, and fact verification, require systems capable of accurately retrieving, validating, and synthesizing information from large-scale knowledge sources. Although Retrieval-Augmented Generation (RAG) has improved the factual grounding of large language models, existing single-agent architectures suffer from incomplete retrieval, inadequate cross-source validation, and limited compositional reasoning. To address these challenges, this paper proposes MARCO (Multi-Agent Collaborative Retrieval-Augmented Generation), a novel framework that decomposes the RAG pipeline into four specialized cooperative agents: a Retriever Agent for hybrid dense-sparse evidence acquisition, a Reasoner Agent for dynamic query decomposition and multi-step inference, a Validator Agent for cross-source consistency assessment and confidence-based filtering, and a Synthesizer Agent for coherent evidence-grounded response generation. Agent interactions are coordinated through a structured Shared Memory Pool and a formal communication protocol. Experimental evaluation on Natural Questions, HotpotQA, and MuSiQue demonstrates that MARCO achieves improvements of 4.2–8.1% in Exact Match and 4.5–7.8% in F1 over strong single-agent RAG baselines, with ablation studies confirming the contribution of each agent module. MARCO establishes a scalable, modular paradigm for next-generation knowledge-intensive AI systems.
No takes yet. Share an insight, caveat, or question.
Satpute et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: