Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 23, 2026MedicineOpen Access

Artificial intelligence in medical education

View Full Paper
Ask AI
Bookmark
Share

Authors

KKKyung Wook KimHJHong Bae JeonJCJong Hyuk Choi

Discussion

Loading...

Member takes

Overview

Comparative evaluation demonstrates superior examination performance by GPT-4 relative to medical students, indicating strong diagnostic capabilities alongside educational utility.

Key Points

  • To evaluate the problem-solving accuracy, reliability, and answer regeneration performance of ChatGPT models (3.5 and 4.0) on medical school examination questions compared with medical students.
  • Administered 49 text-based multiple-choice questions spanning preventive medicine and psychiatry to 41 fourth-year medical students and both ChatGPT models (GPT-3.5 and GPT-4.0).
  • Assessed primary accuracy on the initial prompt and evaluated answer regeneration performance by re-prompting the models following incorrect responses.
  • Fourth-year medical students achieved a mean score of 29.5 ± 5.9 out of 49.
  • Initial scores were 28 for GPT-3.5 and 43 for GPT-4.0, which increased upon re-prompting to 40 and 46, respectively, with GPT-4.0 outperforming all individual students after re-prompting.
  • Both AI models demonstrated significantly higher accuracy than students on easy questions, whereas performance differences between students and the initial GPT-3.5 prompt on difficult questions were nonsignificant.

Cite This Study

Kim et al. (2026) studied this question.

synapsesocial.com/papers/6a8aadb47677a3411444615fhttps://doi.org/10.1097/md.0000000000050323
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models2023 · 3,880 citations
  2. 2Evaluation of Reliability, Repeatability, Robustness, and Confidence of GPT-3.5 and GPT-4 on a Radiology Board–style Examination2024 · 77 citations
  3. 3Comparison of the problem-solving performance of ChatGPT-3.5, ChatGPT-4, Bing Chat, and Bard for the Korean emergency medicine board examination question bank2024 · 19 citations
  4. 4Application of ChatGPT for Orthopedic Surgeries and Patient Care2024 · 43 citations
  5. 5PLOS-LLM: Can and should AI enable a new paradigm of scientific knowledge sharing?2024 · 4 citations