PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 3, 2026Academic Radiology0 citationsOpen Access

Large Language Models in Clinical Decision Support: A Comparative Analysis of Chat-GPT and Breast Radiologists on ACR Appropriateness Criteria

SES. EscobarJRJustin RheeHZH. Zaki

Key Points

  • The study aims to evaluate ChatGPT's performance in selecting imaging modalities compared to breast radiologists and ACR guidelines.
  • Evaluated ten clinical variants from ACR breast imaging criteria.
  • Rated 81 imaging decisions on a scale from 1 to 9 by ChatGPT and radiologists.
  • Analyzed agreement using generalized estimating equations (GEEs) and Bland-Altman plots.
  • Radiologists had the lowest mean bias relative to ACR at 0.2438 (p = 0.489).
  • ChatGPT versions displayed larger mean biases, with GPT-4o showing 2.463 (p < 0.001).
  • Radiologists exhibited smaller bias and variation compared to ChatGPT versions.

Abstract

Rationale and Objectives:This study evaluates the performance of ChatGPT, a large language model (LLM), in selecting appropriate imaging modalities for breast imaging scenarios using the American College of Radiology (ACR) Appropriateness Criteria (AC).We aim to compare the agreement of ChatGPT with the ACR AC to that of breast radiologists at a single institution in selecting appropriate imaging modalities. Methods/Materials:The study utilized ten randomly selected clinical variants from the ACR AC breast imaging category.Outputs were obtained from ChatGPT-3.5,ChatGPT-4, and ChatGPT-4o using the versions available in July 2024.The ChatGPT versions and four breast radiologists rated the appropriateness of 81 imaging decisions on a scale from 1 to 9. For each imaging option within a clinical scenario, the ratings provided by the radiologists and four independent samplings of ChatGPT's responses were aggregated.Agreement between ratings from ChatGPT, radiologists, and the ACR AC was analyzed using generalized estimating equations (GEEs) and Bland-Altman plots to assess consistency and bias.Results: Radiologists had the lowest overall mean bias (0.2438) relative to the ACR (p = 0.489).All versions of ChatGPT had larger mean biases that were significant (GPT-4o: 2.463; GPT-4: 1.7623, GPT-3.5:2.4691, all p < 0.001).All had a slope bias (p < 0.001), but radiologists had the smallest slope bias.In summary, radiologists were closer to the ACR AC and were oftentimes as variable or even less variable as a group than the same ChatGPT version at the same time. Conclusion:ChatGPT shows promise as an AI tool for imaging decision-making, but current versions lack the accuracy, consistency, and reproducibility demonstrated by experienced radiologists.The study underscores the importance for human oversight in clinical applications and the need for further development to improve ChatGPT's and other LLMs' reliability and alignment with established guidelines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Escobar et al. (2026) studied this question.

synapsesocial.com/papers/69cf5f225a333a821460e129https://doi.org/10.1016/j.acra.2026.03.015
Ask AI
Helpful
Bookmark
Share
View Full Paper