PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 2026BMJ Open23 citationsOpen Access

Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit

View Full Paper
NTNicholas B. TillerAMAlessandro R MarconMZMarco Zenone

Key Points

  • The aim is to evaluate the accuracy, referencing, and readability of AI chatbot responses in health and medical fields susceptible to misinformation.
  • Assessed five chatbots: Gemini, DeepSeek, Meta AI, ChatGPT, and Grok.
  • Used 10 health-related questions across five categories: cancer, vaccines, stem cells, nutrition, and athletic performance.
  • Responses rated by experts on a coding matrix as problematic or non-problematic.
  • Citations evaluated for accuracy and completeness; readability scored using the Flesch Reading Ease formula.
  • 49.6% of responses were categorized as problematic.
  • Grok chatbot produced significantly more highly problematic responses compared to random expectations (p=0.038).
  • Performance was best in vaccines and cancer (mean z-scores of -2.57 and -2.12) and worst in nutrition and athletic performance.
  • Median citation completeness was low at 40%; poor reference quality and chatbot hallucinations observed.
  • All readability scores indicated 'Difficult', at a college level.

Abstract

Objectives Artificial intelligence (AI)-driven chatbots have been rapidly adopted across research, education, business, marketing and medicine. Most interactions, however, come from non-experts using chatbots like search engines, including for everyday health and medical queries. Design We conducted an original study to audit chatbot responses in health and medical fields prone to misinformation. Methods Five popular chatbots were assessed: Gemini (Google), DeepSeek (High-Flyer), Meta AI (Meta), ChatGPT (OpenAI) and Grok (xAI). In February 2025, each chatbot was prompted with 10 questions from five categories: cancer, vaccines, stem cells, nutrition and athletic performance. We deployed an adversarial-like framework, using open- and closed-ended prompts designed to strain models toward misinformation or contraindicated advice. Two experts from each category rated responses as ‘non-problematic’, ‘somewhat problematic’ or ‘ highly problematic’ using a coding matrix based on objective, predefined criteria. Citations were scored for accuracy and completeness, and each response was given a Flesch Reading Ease score. Results Nearly half (49.6%) of responses were problematic: 30% somewhat problematic and 19.6% highly problematic. Response quality did not differ significantly among chatbots (p=0.566) but Grok generated significantly more highly problematic responses than would be expected under a random distribution (z-score +2.07, p=0.038). Performance was strongest in vaccines (mean z-score –2.57) and cancer (–2.12), and weakest in stem cells (+1.25), athletic performance (+3.74) and nutrition (+4.35). Chatbot outputs were consistently expressed with confidence and certainty; from 250 total questions, there were only two refusals to answer (0.8%), both from Meta AI. Reference quality was poor, with a median completeness score of 40% (Q1–Q3: 20–67%). Chatbot hallucinations and fabricated citations precluded any chatbot from producing a fully accurate reference list. All readability scores were graded as ‘Difficult’ (30–50), equivalent to college sophomore–senior level. Conclusions The audited chatbots performed poorly when answering questions in misinformation-prone health and medical fields. Continued deployment without public education and oversight risks amplifying misinformation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tiller et al. (2026) studied this question.

synapsesocial.com/papers/69e07dfe2f7e8953b7cbf049https://doi.org/10.1136/bmjopen-2025-112695
Ask AI
Helpful
Bookmark
Share
View Full Paper