PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 5, 2025Uludağ Üniversitesi Tıp Fakültesi Dergisi0 citationsOpen Access

Evaluating the Performance of Large Language Models in Generating Impressions for Radiology Reports

View Full Paper
HKHasan Emin KayaDSDilek SağlamZYZeynep Yazıcı

Key Points

  • The median scores for LLM outputs were between 4 and 5, indicating general success in generating impressions.
  • No statistically significant difference in performance was detected among the three LLM models (p > 0.05).
  • Three LLMs—ChatGPT, Gemini, and Copilot—were assessed based on their ability to summarize radiology reports effectively.
  • Radiologists evaluated the impressions using a Likert scale, considering completeness and clarity.

Abstract

The aim of the study was to evaluate and compare the performance of three popular large language models (LLMs) in generating impressions for radiology reports in Turkish. ChatGPT, Gemini, and Copilot were used to generate impressions for 50 anonymized radiology reports using a “few-shot” prompt. The impressions were scored by three radiologists using a Likert scale, based on whether they included all relevant information from the report, provided an appropriate summary of the report, contained no misleading information, and could be added to the report without modification. Friedman's test was used to evaluate whether there was a difference between the scores of the LLMs. The 50 reports included 32 magnetic resonance examinations, 11 computed tomography examinations, 5 ultrasound examinations, and 2 fluoroscopy examinations. Of these, 15 were neuroradiology studies, 14 were musculoskeletal studies, 13 were abdominal studies, and 8 were thoracic radiology studies. The median scores for the models’ outputs were 4 and 5. This finding indicates that the radiologists generally found the models successful in generating impressions. Furthermore, no statistically significant difference was found among the models in terms of their performance in containing all information, providing an appropriate summary, avoiding misleading information, and being suitable for inclusion in the report without modification (p = 0.607, 0.327, 0.629, 0.089, respectively). In conclusion, ChatGPT, Gemini, and Copilot were found to be successful in generating impressions for radiology reports in Turkish, and no significant difference in performance was detected among the models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kaya et al. (2025) studied this question.

synapsesocial.com/papers/68bb5f266d6d5674bcd03193https://doi.org/10.32708/uutfd.1653680
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations2023 · 479 citations
  2. 2Constructing a Large Language Model to Generate Impressions from Findings in Radiology Reports2024 · 79 citations
  3. 3Quantitative Evaluation of Large Language Models to Streamline Radiology Report Impressions: A Multimodal Retrospective Analysis2024 · 128 citations
  4. 4Accuracy of ChatGPT, Google Bard, and Microsoft Bing for Simplifying Radiology Reports2023 · 104 citations
  5. 5Diagnostic Performance of ChatGPT from Patient History and Imaging Findings on the Diagnosis Please Quizzes2023 · 151 citations