Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 1, 2025Cardiovascular Diagnosis and TherapyOpen Access

GPT-4.0 with informative prompts achieved the highest diagnostic performance (AUC 0.98, 95% CI: 0.97-0.99) and almost perfect agreement with radiologists (AC1=0.93, 95% CI: 0.90-0.96).

View Full Paper
Ask AI
Bookmark
Share

Why the study?

CMR versatility leads to complex and time-consuming interpretation, but the potential of LLMs for automated classification and diagnosis of CMR reports has not been established.

Does automated interpretation using large language models match the diagnostic accuracy of radiologists in classifying CMR reports for MI, DCM, and HCM?

Population

543 CMR cases of consecutive patients with MI, DCM, or HCM

Comparison

Six LLMs with minimal or informative prompts vs radiologists

Design

Retrospective study

Authors

LWLujing WangLPLiang PengYWYixuan Wan

Discussion

Loading...

Member takes

Overview

LLM-assisted CMR report classification shows retrospective feasibility; leaves open prospective accuracy and clinical workflow integration.

Structured PICO

Does automated interpretation using large language models match the diagnostic accuracy of radiologists in classifying CMR reports for MI, DCM, and HCM?

P
Population
543 cardiac magnetic resonance (CMR) reports of consecutive patients from January 2015 to July 2024, including cases of myocardial infarction (n=275), dilated cardiomyopathy (n=120), and hypertrophic cardiomyopathy (n=148).
I
Intervention
Automated classification and diagnosis using six large language models (GPT-3.5, GPT-4.0, Gemini-1.0, Gemini-1.5, PaLM, and LLaMA) provided with minimal or informative prompts.
C
Comparator
Interpretation and diagnosis by radiologists.
O
Outcome
Classification performance evaluated by accuracy (ACC) and balanced accuracy (BAC), and consistency with radiologists evaluated using Gwet's Agreement Coefficient (AC1 value).

Large language models, particularly GPT-4.0 using informative prompts, demonstrate excellent accuracy and high agreement with radiologists for the automated interpretation of cardiac magnetic resonance reports.

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/6a87cffdf8efd6d7e4fb15a0https://doi.org/10.21037/cdt-2025-112
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluating prompt and data perturbation sensitivity in large language models for radiology reports classification2025 · 10 citations
  2. 2A systematic evaluation of open-source large language models for automated extraction of cardiac MRI parameters from unstructured reports2025
  3. 3Comparative Diagnostic Performance of a Multimodal Large Language Model Versus a Dedicated Electrocardiogram AI in Detecting Myocardial Infarction From Electrocardiogram Images: Comparative Study2025 · 8 citations
  4. 4Evaluation of large language models as a diagnostic tool for medical learners and clinicians using advanced prompting techniques2025 · 10 citations
  5. 5Chatting Ain’t Diagnosing: Diagnostic Variability and Fundamental Errors in Multimodal LLM Interpretation in Radiology2026 · 2 citations