PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 202631 citations

Performance of a large language model on the reasoning tasks of a physician.

View Full Paper
PBPeter G. BrodeurTBThomas A BuckleyZKZahir Kanjee

Key Points

  • This study aims to evaluate the performance of a large language model on clinical reasoning tasks compared to human experts.
  • Evaluated the LLM against hundreds of physicians across five experiments on complex clinical cases.
  • Conducted a real-world study in an emergency room comparing AI and human second opinions.
  • Analyzed performance changes across prior AI generations for clinical decision support.
  • The LLM outperformed physician baselines in all evaluated experiments.
  • Notable improvements were seen in the LLM's reasoning capabilities compared to previous AI generations.
  • The study indicates LLMs have surpassed traditional benchmarks in clinical reasoning.

Abstract

More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Brodeur et al. (2026) studied this question.

synapsesocial.com/papers/69f6e5ac8071d4f1bdfc653fhttps://doi.org/10.1126/science.adz4433
Ask AI
Helpful
Bookmark
Share
View Full Paper