PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 2026European Journal of Anaesthesiology0 citations

Comparative performance of GPT-4 models and expert anaesthesiologists in peri-operative antithrombotic management

View Full Paper
APAntonio Pérez-FerrerEDE. Gredilla DíazJSJesús de Vicente Sánchez

Key Points

  • This research aims to compare the performance of language models and anaesthesiologists in antithrombotic management during surgery.
  • Conducted a cross-sectional analytical study in simulated environments across five anaesthesiology departments in Spain.
  • Designed twenty-five hypothetical scenarios focused on anticoagulant or antiplatelet therapy based on guidelines.
  • Five anaesthesiologists and two language models provided responses to these scenarios.
  • The domain-specific model provided more complete responses than anaesthesiologists and the general-purpose model.
  • Accuracy was similar across groups, with the domain-specific model achieving the highest mean accuracy.
  • Clinicians used external resources in 77% of cases, while the domain-specific model had only 1.3% inaccurate answers.

Abstract

BACKGROUND AND OBJECTIVE: Large language models based on transformer architecture are increasingly considered as clinical decision-support tools; however, evidence of their reliability compared with human experts in high-stakes peri-operative contexts remains limited. Peri-operative antithrombotic management requires precise, guideline-concordant decision-making. The objective of this study was to compare the performance of a general-purpose and a domain-specific, transformer-based language model with that of practising anaesthesiologists in peri-operative antithrombotic scenarios. METHODS: A cross-sectional analytical study was conducted in a fully simulated, nonclinical environment across five university-affiliated anaesthesiology departments in Spain. Twenty-five hypothetical peri-operative scenarios involving anticoagulant or antiplatelet therapy were designed according to current evidence-based guidelines. Five anaesthesiologists generated responses to peer-created and self-authored scenarios (125 clinician responses). Two language models independently answered all scenarios. Completeness and accuracy were rated independently by three blinded experts. RESULTS: The domain-specific model generated more complete responses (mean 4.52 ± 0.39) than anaesthesiologists (3.72 ± 0.43) and the general-purpose model (4.25 ± 0.48; P < 0.001). Accuracy did not differ significantly between groups (P = 0.107), with the highest mean accuracy observed for the domain-specific model (4.31 ± 0.57). Inter-rater reliability ranged from 0.717 to 0.885. The domain-specific model produced no incomplete responses and 1.3% inaccurate answers, whereas clinicians used external resources in 77% of cases. CONCLUSIONS: In simulated peri-operative antithrombotic management scenarios, a domain-specific, transformer-based language model generated faster and more complete responses than clinician-generated answers, while accuracy was comparable across groups. These findings support further prospective evaluation of domain-adapted language models before clinical integration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pérez-Ferrer et al. (2026) studied this question.

synapsesocial.com/papers/69fbe2f2164b5133a91a2516https://doi.org/10.1097/eja.0000000000002418
Ask AI
Helpful
Bookmark
Share
View Full Paper