PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026Life0 citationsOpen Access

Artificial Intelligence for Biomedical Diagnostics: Diagnostic Accuracy and Reliability of Multimodal Large Language Models in Electrocardiogram Interpretation

View Full Paper
HSHenrik StellingAKArmin KrausGGGerrit Grieb

Key Result

Multimodal large language models did not achieve clinically reliable ECG interpretation, with overall categorical accuracy ranging from 52.3% to 64.9% compared to expert-consensus ground truth.

Key Points

  • The research aims to assess the diagnostic accuracy and inter-run reliability of multimodal large language models for ECG interpretation.
  • Evaluated five MLLMs on 13 standard 12-lead ECGs across five independent runs per case.
  • Conducted 2275 task-level assessments for six categorical interpretation tasks.
  • Compared model outputs with expert-consensus ground truth and assessed heart rate estimation using mean absolute error.
  • Categorical accuracy ranged from 52.3% to 64.9% across models.
  • QRS duration classification achieved an accuracy of 66.2–90.8%.
  • ST/T-wave assessment showed the lowest performance at 20.0–41.5%.
  • Heart rate mean absolute error ranged from 14.8 to 46.7 bpm.
  • A dissociation between diagnostic accuracy and inter-run reliability was identified.

Study Design

Type

Observational (n=13)

Structured PICO

Do multimodal large language models accurately and reliably interpret standard 12-lead ECGs compared to expert consensus?

P
Population
13 standard 12-lead ECGs
I
Intervention
Five multimodal large language models (ChatGPT-5.3, Gemini 3.1 Pro, Claude Opus 4.6, Grok 4.1, and ERNIE 5.0) evaluated across five independent runs per case
C
Comparator
Expert-consensus ground truth
O
Outcome
Diagnostic accuracy across six categorical interpretation tasks (rhythm, electrical axis, PR/P-wave morphology, QRS duration, ST/T-wave morphology, and QTc interval) and heart rate estimationsurrogate

Current multimodal large language models demonstrate insufficient diagnostic accuracy and reliability for clinical ECG interpretation.

Abstract

The electrocardiogram (ECG) is a central tool in cardiovascular diagnostics, yet interpretation requires expertise and remains subject to variability. Multimodal large language models (MLLMs) have shown emerging capabilities in medical image analysis, but their performance in ECG interpretation remains insufficiently characterized. This study evaluated the diagnostic accuracy and inter-run reliability of five MLLMs across ECG interpretation tasks. Thirteen standard 12-lead ECGs were presented to five models (ChatGPT-5.3, Gemini 3.1 Pro, Claude Opus 4.6, Grok 4.1, and ERNIE 5.0) across five independent runs per case, yielding 2275 task-level assessments. Six categorical interpretation tasks (rhythm, electrical axis, PR/P-wave morphology, QRS duration, ST/T-wave morphology, and QTc interval) were compared with expert-consensus ground truth, while heart rate estimation was evaluated using mean absolute error (MAE). Overall categorical accuracy ranged from 52.3% to 64.9%. QRS duration classification achieved the highest accuracy (66.2–90.8%), whereas ST/T-wave assessment showed the lowest performance (20.0–41.5%). Heart rate MAE ranged from 14.8 to 46.7 bpm. A dissociation between diagnostic accuracy and inter-run reliability was observed across models. These findings indicate that current MLLMs do not achieve clinically reliable ECG interpretation performance and highlight the importance of assessing diagnostic accuracy and inter-run reliability when evaluating artificial intelligence systems in biomedical diagnostics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Stelling et al. (2026) conducted an observational in Electrocardiogram interpretation (n=13). Multimodal large language models (ChatGPT-5.3, Gemini 3.1 Pro, Claude Opus 4.6, Grok 4.1, ERNIE 5.0) vs. Expert-consensus ground truth was evaluated on Diagnostic accuracy across six categorical interpretation tasks and heart rate mean absolute error. Multimodal large language models did not achieve clinically reliable ECG interpretation, with overall categorical accuracy ranging from 52.3% to 64.9% compared to expert-consensus ground truth.

synapsesocial.com/papers/69e320e740886becb6540186https://doi.org/10.3390/life16040681
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ECG-LM: Understanding Electrocardiogram with a Large Language Model2024 · 30 citations
  2. 2Teaching multimodal LLMs to comprehend 12-lead electrocardiographic images2026 · 7 citations
  3. 3High agreement but low Kappa: I. the problems of two paradoxes1990 · 2,981 citations
  4. 4Generative AI Models (2018–2024): Advancements and Applications in Kidney Care2025 · 15 citations
  5. 5Measuring nominal scale agreement among many raters.1971 · 8,754 citations