PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 29, 2026npj Digital Medicine6 citationsOpen Access

AgentClinic: a multimodal benchmark for tool-using clinical AI agents

View Full Paper
SSSamuel SchmidgallRZRojin ZiaeiCHCarl Harris

Key Points

  • This research aims to evaluate clinical AI agents using a multimodal benchmark to improve decision-making accuracy.
  • Introduced AgentClinic for evaluating large language models in clinical scenarios.
  • Utilized simulated clinical environments with patient interactions and multimodal data.
  • Analyzed the use of various tools and biases in agents.
  • Leveraged electronic health records and conducted a clinical reader study.
  • Found diagnostic accuracies can drop below a tenth due to the sequential decision-making format.
  • Claude-3.5 outperformed other models in most scenarios.
  • Llama-3 showed up to 92% relative improvements with an interactive notebook tool.
  • Experiential learning and adaptive retrieval enhanced the agents' performance.

Abstract

Abstract Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering, which does not accurately depict the complex, sequential nature of clinical decision-making. Here, we introduce AgentClinic, a multimodal agent benchmark for evaluating LLMs in simulated clinical environments that include patient interactions, multimodal data collection under incomplete information, and the usage of various tools, resulting in an in-depth evaluation across nine medical specialties and seven languages. We find that solving MedQA problems in the sequential decision-making format of AgentClinic is considerably more challenging, resulting in diagnostic accuracies that can drop to below a tenth of the original accuracy. Overall, we observe that agents sourced from Claude-3.5 outperform other LLM backbones in most settings. Nevertheless, we see stark differences in the LLMs’ ability to make use of tools, such as experiential learning, adaptive retrieval, and reflection cycles. Strikingly, Llama-3 shows up to 92% relative improvements with the notebook tool that allows for writing and editing notes that persist across cases. To further scrutinize our clinical simulations, we leverage real-world electronic health records, perform a clinical reader study, perturb agents with biases, and explore patient-centric metrics that this interactive environment enables.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Schmidgall et al. (2026) studied this question.

synapsesocial.com/papers/69f154f9879cb923c4945631https://doi.org/10.1038/s41746-026-02674-7
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments2024 · 22 citations
  2. 2Benchmarking large language model-based agent systems for clinical decision tasks2026 · 13 citations
  3. 3MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents2025 · 51 citations
  4. 4Evaluating Large Language Model Diagnostic Performance on JAMA Clinical Challenges via a Multi-Agent Conversational Framework2025
  5. 5Evaluating large language models as agents in the clinic2024 · 136 citations