PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 2026Nature Medicine50 citationsOpen Access

LLM-assisted systematic review of large language models in clinical medicine

SCSully F. ChenAAAnton AlyakinASAndreas Seas

Key Points

  • To assess the utilization and evidence of large language models (LLMs) in clinical medicine.
  • Conducted LLM-assisted systematic review of 4,609 peer-reviewed studies published from January 2022 to September 2025.
  • Analyzed studies based on the use of real-world patient data and trial types.
  • Categorized tasks for LLMs in clinical settings, including communication and knowledge retrieval.
  • Only 1,048 studies utilized real-world patient data, with just 19 being prospective randomized trials.
  • LLMs outperformed humans in 33% of head-to-head comparisons, depending on task realism and training.
  • Most studies (1,857) addressed simulated scenarios or exam-style tasks, with 25% having sample sizes less than 30.

Abstract

Clinical evaluations of large language models (LLMs) have rapidly expanded since 2022, yet their evidence base remains opaque. The overwhelming volume of studies creates challenges for manual curation and review. However, LLMs themselves offer the scalability and capability to evaluate the ever-growing evidence base. This LLM-assisted review identified 4,609 peer-reviewed studies in clinical medicine between January 2022 and September 2025, equating to roughly 3.2 papers per day. Only 1,048 studies used real-world patient data and of these only 19 were prospective randomized trials; most addressed simulated scenarios (n = 1,857) or exam-style tasks (n = 1,704). ChatGPT and related OpenAI models constitute 65.7% of evaluated models, with Gemini/Bard a distant second constituting 13.1% of evaluated models. Patient-facing communication and education comprised 17% of tasks, followed by knowledge retrieval, and education and assessment simulation. Across 1,046 head-to-head comparisons, LLMs outperformed humans in 33% of comparisons, with a strong dependency on task realism and level of training. At least 25% of studies had sample sizes less than 30. Despite the growth of LLMs in medicine, rigorous, patient-centered evidence remains scarce, underscoring the need for larger prospective trials before clinical adoption. A large language model (LLM)-powered systematic review of over 1,000 studies revealed that, despite the growth of medical research involving LLMs, a majority of studies do not involve real-world clinical data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2026) studied this question.

synapsesocial.com/papers/69a91e65d6127c7a504c25c1https://doi.org/10.1038/s41591-026-04229-5
Ask AI
Helpful
Bookmark
Share
View Full Paper