Why the study?
CMR versatility leads to complex and time-consuming interpretation, but the potential of LLMs for automated classification and diagnosis of CMR reports has not been established.
Does automated interpretation using large language models match the diagnostic accuracy of radiologists in classifying CMR reports for MI, DCM, and HCM?
Population
543 CMR cases of consecutive patients with MI, DCM, or HCM
Comparison
Six LLMs with minimal or informative prompts vs radiologists
Design
Retrospective study
Authors
Loading...
LLM-assisted CMR report classification shows retrospective feasibility; leaves open prospective accuracy and clinical workflow integration.
Does automated interpretation using large language models match the diagnostic accuracy of radiologists in classifying CMR reports for MI, DCM, and HCM?
Large language models, particularly GPT-4.0 using informative prompts, demonstrate excellent accuracy and high agreement with radiologists for the automated interpretation of cardiac magnetic resonance reports.
Wang et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: