PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 13, 2025npj Digital Medicine339 citationsOpen Access

A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation

View Full Paper
EAElham AsgariSt Thomas' HospitalNMNina Montaña-BrownNicolaus Copernicus UniversityMDMagda DuboisWellcome Centre for Human Neuroimaging

Key Points

Key points are not available for this paper at this time.

Abstract

Integrating large language models (LLMs) into healthcare can enhance workflow efficiency and patient care by automating tasks such as summarising consultations. However, the fidelity between LLM outputs and ground truth information is vital to prevent miscommunication that could lead to compromise in patient safety. We propose a framework comprising (1) an error taxonomy for classifying LLM outputs, (2) an experimental structure for iterative comparisons in our LLM document generation pipeline, (3) a clinical safety framework to evaluate the harms of errors, and (4) a graphical user interface, CREOLA, to facilitate these processes. Our clinical error metrics were derived from 18 experimental configurations involving LLMs for clinical note generation, consisting of 12,999 clinician-annotated sentences. We observed a 1.47% hallucination rate and a 3.45% omission rate. By refining prompts and workflows, we successfully reduced major errors below previously reported human note-taking rates, highlighting the framework's potential for safer clinical documentation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Asgari et al. (2025) studied this question.

synapsesocial.com/papers/6968b68386a1f8fec068321ehttps://doi.org/10.1038/s41746-025-01670-7
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Using ChatGPT to write patient clinic letters2023 · 402 citations
  2. 2Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation2024 · 389 citations
  3. 3Creating Trustworthy LLMs: Dealing with Hallucinations in Healthcare AI2023 · 20 citations
  4. 4medIKAL: Integrating Knowledge Graphs as Assistants of LLMs for Enhanced Clinical Diagnosis on EMRs2024 · 4 citations
  5. 5Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena2023 · 485 citations