PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025npj Digital Medicine53 citationsOpen Access

Accelerating clinical evidence synthesis with large language models

View Full Paper
ZWZifeng WangShangrao Normal UniversityLCLang CaoUniversity of Illinois Urbana-ChampaignBDBenjamin DanekNational Institutes of Health

Key Points

  • Human-AI collaboration with TrialMind improved recall by 71.4% and reduced screening time by 44.2%.
  • For data extraction, TrialMind outperformed GPT-4 by 16-32% in accuracy, highlighting the efficiency of new tools.
  • The study revealed TrialMind’s synthesized evidence was preferred by medical experts in 62.5%-100% of cases over traditional methods.
  • Using TrialMind, high recall rates of 0.711-0.834 were achieved compared to much lower human baseline rates of 0.138-0.232.

Abstract

Clinical evidence synthesis largely relies on systematic reviews (SR) of clinical studies from medical literature. Here, we propose a generative artificial intelligence (AI) pipeline named TrialMind to streamline study search, study screening, and data extraction tasks in SR. We chose published SRs to build TrialReviewBench, which contains 100 SRs and 2,220 clinical studies. For study search, it achieves high recall rates (Ours 0.711-0.834 v.s. Human baseline 0.138-0.232). For study screening, TrialMind beats previous document ranking methods in a 1.5-2.6 fold change. For data extraction, it outperforms a GPT-4's accuracy by 16-32%. In a pilot study, human-AI collaboration with TrialMind improved recall by 71.4% and reduced screening time by 44.2%, while in data extraction, accuracy increased by 23.5% with a 63.4% time reduction. Medical experts preferred TrialMind's synthesized evidence over GPT-4's in 62.5%-100% of cases. These findings show the promise of accelerating clinical evidence synthesis driven by human-AI collaboration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68c1bd2a54b1d3bfb60ee131https://doi.org/10.1038/s41746-025-01840-7
Ask AI
Helpful
Bookmark
Share
View Full Paper