PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 26, 202538 citationsOpen Access

AI Agents in Clinical Medicine: A Systematic Review

View Full Paper
AGAlon GorenshteinMOMahmud OmarBGBenjamin S. Glicksberg

Key Points

  • AI agents improved performance in clinical tasks significantly, with some showing over 60 percentage points increase compared to standard models.
  • The review included twenty studies that demonstrated superior accuracy in AI agent systems versus baseline large language models in various tasks.
  • AI agents effectively handled discrete tasks like medication dosing and evidence retrieval, particularly performing well in complex scenarios.
  • Future research calls for prospective trials that use real-world data to validate safety and cost-effectiveness of these AI systems.

Abstract

Background: AI agents built on large language models (LLMs) can plan tasks, use external tools, and coordinate with other agents. Unlike standard LLMs, agents can execute multi-step processes, access real-time clinical information, and integrate multiple data sources. There has been interest in using such agents for clinical and administrative tasks, however, there is limited knowledge on their performance and whether multi-agent systems function better than a single agent for healthcare tasks. Purpose: To evaluate the performance of AI agents in healthcare, compare AI agent systems vs. standard LLMs and catalog the tools used for task completion Data Sources: PubMed, Web of Science, and Scopus from October 1, 2022, through August 5, 2025. Study Selection: Peer-reviewed studies implementing AI agents for clinical tasks with quantitative performance comparisons. Data Extraction: Two reviewers (A.G., M.O.) independently extracted data on architectures, performance metrics, and clinical applications. Discrepancies were resolved by discussion, with a third reviewer (E.K.) consulted when consensus could not be reached. Data Synthesis: Twenty studies met inclusion criteria. Across studies, all agent systems outperformed their baseline LLMs in accuracy performance. Improvements ranged from small gains to increases of over 60 percentage points, with a median improvement of 53 percentage points in single-agent tool-calling studies. These systems were particularly effective for discrete tasks such as medication dosing and evidence retrieval. Multi-agent systems showed optimal performance with up to 5 agents, and their effectiveness was particularly pronounced when dealing with highly complex tasks. The highest performance boost occurred when the complexity of the AI agent framework aligned with that of the task. Limitations: Heterogeneous outcomes precluded quantitative meta-analysis. Several studies relied on synthetic data, limiting generalizability. Conclusions: AI agents consistently improve clinical task performance of Base-LLMs when architecture matches task complexity. Our analysis indicates a step-change over base-LLMs, with AI agents opening previously inaccessible domains. Future efforts should be based on prospective, multi-center trials using real-world data to determine safety, task matched and cost-effectiveness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gorenshtein et al. (2025) studied this question.

synapsesocial.com/papers/68af620aad7bf08b1eae313fhttps://doi.org/10.1101/2025.08.22.25334232
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Artificial intelligence agents in healthcare research: A scoping review2026 · 19 citations
  2. 2Benchmarking large language model-based agent systems for clinical decision tasks2026 · 13 citations
  3. 3Autonomous Artificial Intelligence Agents for Clinical Decision Making in Oncology2024 · 5 citations
  4. 4A Review of Multi-Agent AI Systems for Biological and Clinical Data Analysis2026 · 1 citations
  5. 5Muli-Agent AI Systems in Healthcare: Technical and Clinical Analysis2024 · 7 citations