Randomized trial demonstrates the importance of careful AI integration in evidence synthesis workflows, suggesting responsible use and evaluation are crucial.
Evidence synthesis forms the foundation of evidence-based medicine by systematically reviewing and integrating findings from multiple studies to inform clinical decisions, guidelines, and policy. This process is labor-intensive and often requires multidisciplinary expertise. The growing accessibility of generative artificial intelligence (AI) tools has led to their integration into evidence synthesis workflows.1 However, how we integrate these tools responsibly – balancing efficiency while maintaining rigor – and how we evaluate that balance, is still being determined. Long et al.2 (a self-described multidisciplinary team of novice AI users) offer a model for conservative integration. They documented their process, including prompts and failure modes, and published a standard operating procedure (SOP) alongside their case study. The operational detail of the SOP alone is a meaningful contribution. The substantive value of the work lies in how cognitive labor was distributed. AI tools were only integrated to perform bounded tasks; inclusion decisions and data extraction remained human-led. Cognitive offloading, or shifting decision making from humans to generative AI, is difficult to do responsibly. Even experienced AI users must approach these tools with caution to ensure that the resulting process maintains its rigor. The discipline with which the authors constrained the technology's role is what makes the case study useful guidance for similar multidisciplinary teams learning the technical side as they go. Joint guidance from Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence reinforces this ‘human-in-the-loop’ approach to AI integration.3 The case study does not compare its workflow against a non-AI baseline (nor does it claim to). Thus, claims about efficiency are based on team experience rather than a quantitative comparison. The methodology also relies on commercial AI platforms; no SOP can fully mitigate reproducibility issues arising from technologies whose behavior can change between releases, as the authors themselves report. Long et al.'s contribution should therefore be read as a carefully documented operational case, not as evidence that generative AI improved the rapid evidence mapping reported in this work. Case reports of AI integration in evidence synthesis are becoming quite common, and justifiably so. They provide value to researchers exploring AI integration within their own work, especially when methods are transparently described as they are in this work. However, we should also be thinking about how we move beyond describing how AI was used to focus on evaluating what its use changed. When considering AI integration into evidence synthesis workflows, we should be asking ourselves what this integration is expected to improve and what level of performance would be considered acceptable. For example, screening decisions made with AI support could be (and have been)4 compared with human-only screening, and total workflow time could be compared between fully human workflows and those with AI integration (the latter of which should include time for human verification). Trade-offs should be explicitly described. A workflow that provides time savings while missing relevant studies or reducing reproducibility may not actually represent an improvement. However, a tool that introduces predictable errors while substantially improving efficiency may be useful if those trade-offs are justified against the project aims. Without formal evaluation mechanisms, we have little basis to determine whether case studies showing AI integration meaningfully improve the workflows they describe. Not required.
No takes yet. Share an insight, caveat, or question.
Kristen L. Scotti (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: