PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 25, 2026npj Digital Medicine6 citationsOpen Access

The effect of medical explanations from large language models on diagnostic accuracy in radiology

PSPhilipp SpitzerDHDaniel HendriksJRJan Rudolph

Key Points

  • This study aims to determine how different formats of explanations from large language models affect diagnostic accuracy among radiologists.
  • Conducted a randomized experiment with N = 2020 assessments
  • Participants received either no support or one of three LLM-generated explanation formats: standard output, differential diagnosis, or chain-of-thought reasoning
  • Evaluated the impact of explanation formats on diagnostic accuracy across varying case difficulties and radiologist backgrounds.
  • Chain-of-thought explanations improved diagnostic accuracy by 12.2% compared to the control (P = 0.001)
  • Chain-of-thought outperformed standard output by 7.2% (P = 0.040) and differential diagnosis by 9.7% (P = 0.004)
  • Results indicate that well-explained reasoning increases diagnostic performance and helps correct potential LLM prediction errors.

Abstract

Abstract Large language models (LLMs) are increasingly used by physicians for diagnostic support. A key advantage of LLMs is the ability to generate explanations that can help physicians understand the reasoning behind a diagnosis. However, the best-suited format for LLM-generated explanations remains unclear. In this large-scale study, we examined the effect of different formats for LLM explanations on clinical decision-making. For this, we conducted a randomized experiment with radiologists reviewing patient cases with radiological images ( N = 2020 assessments). Participants received either no LLM support (control group) or were supported by one of three LLM-generated explanations: (1) a standard output providing the diagnosis without explanation; (2) a differential diagnosis comparing multiple possible diagnoses; or (3) a chain-of-thought explanation offering a detailed reasoning process for the diagnosis. We find that the format of explanations significantly influences diagnostic accuracy. The chain-of-thought explanations yielded the best performance, improving the diagnostic accuracy by 12.2% compared to the control condition without LLM support ( P = 0.001). The chain-of-thought explanations are also superior to the standard output without explanation ( + 7.2%; P = 0.040) and the differential diagnosis format ( + 9.7%; P = 0.004). We further assessed the robustness of these findings across case difficulty and different physician backgrounds, such as general vs. specialized radiologists. Evidently, in the controlled setting of our vignette study, explaining the reasoning for a diagnosis helps physicians to identify and correct potential errors in LLM predictions and thus improve overall decisions. Altogether, the results highlight the importance of explanations in medical LLMs to support the reasoning processes of physicians, so that medical LLMs can improve diagnostic performance and, ultimately, patient outcomes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Spitzer et al. (2026) studied this question.

synapsesocial.com/papers/69ec5b6088ba6daa22dace55https://doi.org/10.1038/s41746-026-02619-0
Ask AI
Helpful
Bookmark
Share
View Full Paper