PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 26, 2026BMJ Health & Care Informatics6 citationsOpen Access

Omission and hallucination prevalence of clinical guidelines in diagnostic large language model outputs

View Full Paper
RKRobin van KesselMAMichael AndersonBMBrian; id_orcid 0000-0002-0683-3877 McMillan

Key Points

  • This research assesses how large language models incorporate clinical guidelines, focusing on omissions and hallucinations across various conditions.
  • Simulated case vignettes were used to generate outputs from GPT-4.1 and DeepSeek-V3.
  • Evaluations were conducted across cases of hypercholesterolaemia and type-2 diabetes mellitus with varying sociodemographic characteristics.
  • A total of 12,197 outputs were analyzed for omissions and hallucinations based on pre-defined criteria.
  • Omissions reached up to 97% for DeepSeek-V3 and 46% for GPT-4.1 across outputs.
  • Hallucinations occurred in up to 9% of outputs, with guideline citations ranging from 0% to 78.39%.
  • Variation in output quality was primarily influenced by patient location rather than sex or ethnicity.

Abstract

Objective Meaningful assessments of how large language models (LLMs) incorporate clinical guidelines require large-scale testing over many queries. Here, we evaluate the prevalence of clinical guideline omissions and hallucinations in a large sample of diagnostic LLM outputs. Methods We used simulated case vignettes and zero-shot prompting to generate diagnostic outputs and rationales from GPT-4.1 and DeepSeek-V3. English case vignettes were created for hypercholesterolaemia and type-2 diabetes mellitus. Each vignette contained identical medical information, while sociodemographic characteristics varied in terms of sex, ethnicity and location. We calculated the prevalence of existing and hallucinated clinical guidelines in LLM outputs across disease, LLM and sociodemographic characteristics. Results We analysed a total of 12 197 LLM outputs, which quantifies three hazard areas: omissions (up to 97% for DeepSeek-V3 and 46% for GPT-4.1), hallucinations (up to 9%) and inconsistencies (guideline citation rate ranging from 0% to 78.39% across sociodemographic vignettes). Omission and hallucination rates were generally similar across vignettes with different sex or ethnicity data, yet were particularly sensitive to patient location. Discussion This study highlights significant variability in clinical guideline prediction across two different diseases, three different sociodemographic variables and two LLMs, even when the LLMs were instructed by identical prompts, establishing clinical guideline prediction in LLM outputs as a stochastic event. Conclusion The stochastic nature of LLMs creates a unique challenge for evidence generation and clinical deployment. Being able to measure and capture this stochasticity within high-quality research designs will be a prerequisite to advancing the responsible deployment of LLMs in healthcare.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kessel et al. (2026) studied this question.

synapsesocial.com/papers/69edae394a46254e215b57fehttps://doi.org/10.1136/bmjhci-2025-101959
Ask AI
Helpful
Bookmark
Share
View Full Paper