PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2024Research Synthesis Methods230 citationsOpen Access

Can large language models replace humans in systematic reviews? Evaluating GPT‐4's efficacy in screening and extracting data from peer‐reviewed and grey literature in multiple languages

View Full Paper
QKQusai KhraishaSPS. Van PutJKJohanna Kappenberg

Key Points

Key points are not available for this paper at this time.

Abstract

Systematic reviews are vital for guiding practice, research and policy, although they are often slow and labour-intensive. Large language models (LLMs) could speed up and automate systematic reviews, but their performance in such tasks has yet to be comprehensively evaluated against humans, and no study has tested Generative Pre-Trained Transformer (GPT)-4, the biggest LLM so far. This pre-registered study uses a "human-out-of-the-loop" approach to evaluate GPT-4's capability in title/abstract screening, full-text review and data extraction across various literature types and languages. Although GPT-4 had accuracy on par with human performance in some tasks, results were skewed by chance agreement and dataset imbalance. Adjusting for these caused performance scores to drop across all stages: for data extraction, performance was moderate, and for screening, it ranged from none in highly balanced literature datasets (~1:1) to moderate in those datasets where the ratio of inclusion to exclusion in studies was imbalanced (~1:3). When screening full-text literature using highly reliable prompts, GPT-4's performance was more robust, reaching "human-like" levels. Although our findings indicate that, currently, substantial caution should be exercised if LLMs are being used to conduct systematic reviews, they also offer preliminary evidence that, for certain review tasks delivered under specific conditions, LLMs can rival human performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khraisha et al. (2024) studied this question.

synapsesocial.com/papers/68e73fdcb6db6435876b94e4https://doi.org/10.1002/jrsm.1715
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Cardiovascular mortality and oral antidiabetic drugs: protocol for a systematic review and network meta- analysis2014 · 2,234 citations
  2. 2Breaking bad news in the emergency department: a randomized controlled study of a training using role-play simulation2018 · 347 citations
  3. 3Library Hi Tech2023 · 9 citations
  4. 4Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry2017 · 767 citations
  5. 5Making progress with the automation of systematic reviews: principles of the International Collaboration for the Automation of Systematic Reviews (ICASR)2018 · 201 citations