PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 2025Medical Education3 citations

‘ChatGPT can make mistakes’ warnings fail: A randomized controlled trial

View Full Paper
YKYavuz Selim KıyakÖÇÖzlem ÇoşkunIBIşıl İrem Budakoğlu

Key Points

  • Warnings did not alter diagnostic changes among medical students, with change rates of 15.3% in no-warning and 15.9% in warning groups.
  • The weight-of-advice was significantly lower than average, indicating underweighting of AI diagnostic advice among students.
  • Mixed-effects models were utilized to analyze intervention effects in a randomized controlled trial with 186 medical students.
  • The findings refine advice-taking theory and suggest simple warnings may not ensure calibrated trust in AI systems.

Abstract

Abstract Background Warnings are commonly used to signal the fallibility of AI systems like ChatGPT in clinical decision‐making. Yet, little is known about whether such disclaimers influence medical students' diagnostic behaviour. Drawing on the Judge–Advisor System (JAS) theory, we investigated whether the warning alters advice‐taking behaviour by modifying perceived advisor credibility. Method In this randomized controlled trial, 186 fourth‐year medical students evaluated three clinical vignettes with two diagnostic options. Each case was specifically designed to include the presentations of both diagnoses to make the case ambiguous. Students were randomly assigned to receive feedback either with (warning arm) or without (no‐warning arm) a prominently displayed warning (‘ChatGPT can make mistakes. Check important info’.). After submitting their initial response, students received ChatGPT‐attributed disagreeing diagnostic feedback explaining why the alternate diagnosis was correct. Then they were given the opportunity to revise their original choice. Advice‐taking was measured by whether students changed their diagnosis after viewing AI input. We analysed change rates, weight‐of‐advice (WoA) and used mixed‐effects models to assess intervention effects. Results The warning did not influence diagnostic changes (15.3% no‐warning vs. 15.9% warning; OR = 1.09, 95% CI: 0.46–2.59, p = 0.84). The WoA was 0.15 (SD = 0.36), significantly lower than the 0.30 average in prior JAS meta‐analysis ( p < 0.001). Among students who retained their original diagnosis, the warning group showed a tendency toward providing explanations on why they disagree with the AI advisor (60% vs. 51%, p = 0.059). Conclusions The students underweight AI's diagnostic advice. The disclaimer did not alter students' use of AI advice, suggesting that their perceived credibility of ChatGPT was already near a behavioural floor. This finding supports the existence of a credibility threshold, beyond which additional cautionary cues have limited effect. Our results refine advice‐taking theory and signal that simple warnings may be insufficient to ensure calibrated trust in AI‐supported learning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kıyak et al. (2025) studied this question.

synapsesocial.com/papers/68d9052941e1c178a14f5877https://doi.org/10.1111/medu.70056
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis2024 · 134 citations
  2. 2Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences2006 · 1,027 citations
  3. 3Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems2025 · 532 citations
  4. 4Automation bias: a systematic review of frequency, effect mediators, and mitigators2011 · 1,046 citations
  5. 5“Always check important information!” - The role of disclaimers in the perception of AI-generated content2025 · 25 citations