PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 29, 2026Frontiers in Digital Health0 citationsOpen Access

Blinded two-phase evaluation of large language models in complex cardiac surgery: task-specific performance and human-AI collaboration

MLMarc LéonStanford UniversityRFR FengStanford UniversityMFManuel Quiroz FloresStanford University

Key Points

Key points are not available for this paper at this time.

Abstract

Background Large language models (LLMs) have demonstrated strong performance on standardized medical benchmarks. However, their potential in complex surgical decision-making is largely uncharacterized. Critically, human–LLM collaboration regarding the extent to which clinicians can effectively recognize and integrate model-generated reasoning has emerged as an unaddressed question. To address these gaps, we developed a two-phase evaluation framework to simultaneously assess LLM performance and human–LLM collaboration in cardiac surgery. Methods A panel of senior cardiac surgeons independently developed 15 high-fidelity cardiac surgery scenarios, each paired with a clinically relevant open-ended reasoning task, expert-curated reference answers, and a 10-dimensional weighted evaluation framework. Five representative LLMs (O1, O3-mini-high, DeepSeek-R1, GPT-4, and Llama3-OpenBioLLM-70B) were prompted using a multi-agent strategy. A separate group of senior surgeons conducted a blinded two-phase evaluation to assess model performance and evaluator judgment shifts: in the first round, they rated LLMs independently; in the second, they were shown the reference answers and invited to revise their ratings, with changes being optional. Results LLM performance varied across scenarios, but relative rankings remained stable. Median normalized scores were highest for O1 (0.896), followed by O3-mini-high (0.854), DeepSeek-R1 (0.792), GPT-4 (0.667), and Llama3-OpenBioLLM-70B (0.521). Across evaluation dimensions, scenario comprehension scored highest (0.920), while patient safety (0.507), hallucination avoidance (0.549), and clinical efficiency (0.597) were lowest across models. Second-round normalized scores declined for four LLMs, with 7.57% of ratings revised from affirmative to negative and only 2.59% from negative to affirmative. Among the five highest-weighted evaluation dimensions, 10.16% of second-round ratings were revised from affirmative to negative. Conclusions Reasoning-optimized LLMs outperformed all other models. However, all models exhibited clinical limitations, including poor performance in core evaluation dimensions and scenarios requiring complex, longitudinal reasoning tasks. Overacceptance was the dominant collaboration imbalance, reflecting that clinicians over-accepted model reasoning that appears clinically sound yet is incorrect or potentially harmful. These findings suggest that these LLMs are not yet ready for safe use in complex surgical settings due to both performance limitations and human–LLM collaboration imbalance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Léon et al. (2026) studied this question.

synapsesocial.com/papers/6a2364a5a9ddd97f33f1f041https://doi.org/10.3389/fdgth.2026.1769467
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1An evaluation framework for clinical use of large language models in patient interaction tasks2025 · 182 citations
  2. 2Can ChatGPT transform cardiac surgery and heart transplantation?2024 · 19 citations
  3. 3The diagnostic and triage accuracy of the GPT-3 artificial intelligence model: an observational study2024 · 102 citations
  4. 4Large Language Model Influence on Diagnostic Reasoning2024 · 649 citations
  5. 5Almanac — Retrieval-Augmented Language Models for Clinical Medicine2024 · 390 citations