PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 3, 2026Digital Health0 citationsOpen Access

ChatGPT response consistency to the 2025 ESC/EACTS guidelines for the management of valvular heart disease: A test–retest study using binary and multiple-choice questions

View Full Paper
ÇMÇetin MirzaoğluZUZeynep UlutaşYKYücel Karaca

Key Points

  • This study aims to assess the response consistency of ChatGPT-5.2 to clinical questions based on specific guidelines.
  • Prospective observational study using a test–retest design.
  • Administered 100 guideline-based questions (60 binary, 40 multiple-choice) to ChatGPT-5.2 over two occasions with a 14-day interval.
  • Responses were evaluated by two cardiologists and analyzed using McNemar's test and Cohen’s kappa coefficient.
  • Binary questions showed consistent accuracy of 96.7% in both assessments.
  • Multiple-choice accuracy improved from 75.0% to 87.5%, and overall accuracy rose from 88.0% to 93.0%.
  • Cohen’s kappa indicated moderate agreement for binary questions and low for multiple-choice questions.

Abstract

Background/Objectives This study aimed to evaluate the response variability and temporal instability of responses generated by the artificial intelligence–based model ChatGPT-5.2 to structured clinical questions derived from the 2025 ESC/EACTS Guidelines for the Management of Valvular Heart Disease (GMVHD). Methods This prospective observational study employed a test–retest design. A structured set of 100 guideline-based questions—comprising 60 binary (true/false) and 40 multiple-choice items—was developed by two cardiologists (Ç.M. and Z.U.) in accordance with the 2025 ESC/EACTS GMVHD. The question set was administered to ChatGPT-5.2 on two separate occasions with a 14-day interval. The model was instructed to provide answers only, without any explanatory commentary. ChatGPT-generated responses were independently evaluated and coded as correct or incorrect by two cardiologists (Y.K. and Ç.M.). Numerical changes in responses were assessed using McNemar’s test, while test–retest reliability was evaluated using Cohen’s kappa coefficient. Results For binary questions, ChatGPT demonstrated an accuracy of 96.7% in both assessments. Accuracy for multiple-choice questions increased from 75.0% at baseline to 87.5% at the second assessment. When all questions were analyzed together, overall accuracy improved from 88.0% to 93.0%. A numerical increase in accuracy was observed between T1 and T2, without a statistically significant temporal difference. (p > 0.05). Cohen’s kappa analysis indicated moderate agreement for binary questions and low agreement for multiple-choice questions. Conclusion ChatGPT-5.2 demonstrated numerical improvement without statistically significant temporal difference and short-term performance change when answering guideline-based clinical questions on valvular heart disease. However, the relatively high initial error rate in multiple-choice questions represents a limitation for clinical reliability. At present, AI systems may be considered supportive tools for guideline-based information retrieval and clinical education.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mirzaoğlu et al. (2026) studied this question.

synapsesocial.com/papers/6a1fc550dee9eb8c0dce6c51https://doi.org/10.1177/20552076261458145
Ask AI
Helpful
Bookmark
Share
View Full Paper