PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025Communications Medicine155 citationsOpen Access

Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support

View Full Paper
MOMahmud OmarVSVera SorinJCJeremy D. Collins

Key Points

  • Hallucination rates in large language models ranged from 50% to 82%, indicating significant vulnerability.
  • Using a mitigating prompt reduced hallucination rates from 66% to 44%, showing promise in error correction.
  • Testing involved 300 physician-validated vignettes, assessing responses across six LLMs under different conditions.
  • Temperature adjustments did not significantly improve hallucination rates, indicating that alternative strategies are needed.

Abstract

Large language models (LLMs) show promise in clinical contexts but can generate false facts (often referred to as "hallucinations"). One subset of these errors arises from adversarial attacks, in which fabricated details embedded in prompts lead the model to produce or elaborate on the false information. We embedded fabricated content in clinical prompts to elicit adversarial hallucination attacks in multiple large language models. We quantified how often they elaborated on false details and tested whether a specialized mitigation prompt or altered temperature settings reduced errors. We created 300 physician-validated simulated vignettes, each containing one fabricated detail (a laboratory test, a physical or radiological sign, or a medical condition). Each vignette was presented in short and long versions-differing only in word count but identical in medical content. We tested six LLMs under three conditions: default (standard settings), mitigating prompt (designed to reduce hallucinations), and temperature 0 (deterministic output with maximum response certainty), generating 5,400 outputs. If a model elaborated on the fabricated detail, the case was classified as a "hallucination". Hallucination rates range from 50 % to 82 % across models and prompting methods. Prompt-based mitigation lowers the overall hallucination rate (mean across all models) from 66 % to 44 % (p < 0.001). For the best-performing model, GPT-4o, rates decline from 53 % to 23 % (p < 0.001). Temperature adjustments offer no significant improvement. Short vignettes show slightly higher odds of hallucination. LLMs are highly susceptible to adversarial hallucination attacks, frequently generating false clinical details that pose risks when used without safeguards. While prompt engineering reduces errors, it does not eliminate them.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Omar et al. (2025) studied this question.

synapsesocial.com/papers/68c1aabf54b1d3bfb60e2e8ehttps://doi.org/10.1038/s43856-025-01021-3
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis2024 · 380 citations
  2. 2Use of Generative AI to Identify Helmet Status Among Patients With Micromobility-Related Injuries From Unstructured Clinical Notes2024 · 30 citations
  3. 3Detecting hallucinations in large language models using semantic entropy2024 · 810 citations
  4. 4The Role of Prompt Engineering for Multimodal LLM Glaucoma Diagnosis2024 · 13 citations
  5. 5A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions2023 · 218 citations