PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 2026Turkish Journal of Emergency Medicine0 citationsOpen Access

Evaluation of ChatGPT’s performance on emergency medicine board examination questions

View Full Paper
MGMustafa Can GüzelceSÖS. ÖzgürİŞİlker Şalli

Key Points

  • This research aims to assess ChatGPT's ability to correctly answer emergency medicine board examination questions.
  • Conducted a cross-sectional observational study using 25 standardized questions from the Turkish Board of Emergency Medicine.
  • Manually input questions into two models: GPT-4 and GPT-4o using OpenAI interface.
  • Evaluated model responses for accuracy and consistency, focusing on domain-specific errors.
  • GPT-4 answered 80% of questions accurately on the first attempt, improving to 84% on repetition.
  • GPT-4o achieved an 88% accuracy rate on its first attempt, showing consistency in repeated assessments.
  • Identified errors in specific domains: trauma during pregnancy, pediatric resuscitation, and adult resuscitation.

Abstract

Abstract: OBJECTIVES: We aimed to evaluate the performance of a large language model (ChatGPT) in answering official sample questions from the Turkish Board of Emergency Medicine (TBEM). Two versions of the model, GPT-4 and GPT-4o, were assessed to explore consistency and accuracy across iterations. METHODS: A cross-sectional observational study was conducted using 25 standardized multiple-choice questions publicly released by TBEM. Each question was manually entered into GPT-4 and GPT-4o through the OpenAI interface. Both models were prompted to select the best single answer from the provided options without additional clarification or training context. Model responses were evaluated for accuracy, consistency upon repetition, and domain-specific error types. This study is compliant with the STROBE statement and the MedinAI reporting guidelines. RESULTS: GPT-4 correctly answered 20 out of 25 questions (80%) on the first attempt. On repetition, its score improved to 84%. GPT-4o also achieved a score of 88% (22/25) on its first attempt and showed consistent results upon a second evaluation, providing identical answers in both trials. Errors occurred in the domains of trauma during pregnancy, pediatric resuscitation, and adult resuscitation protocols. Both models demonstrated strong performance in fact-based domains and in questions involving image descriptions. CONCLUSION: GPT-4 and GPT-4o performed above the TBEM passing threshold, showing solid accuracy across a range of emergency medicine topics. Both excelled in fact-based and image-related questions. However, they showed limitations in clinical reasoning, particularly in scenarios requiring nuanced judgment. These tools may support examination preparation but should not replace the expertise of trained emergency physicians.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Güzelce et al. (2026) studied this question.

synapsesocial.com/papers/69d1fde4a79560c99a0a445ahttps://doi.org/10.4103/tjem.tjem_262_25
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The Role of Large Language Models in Pediatric Emergency Medicine: Accuracy and Decision-Support Potential of ChatGPT2026
  2. 2Generative AI’s performance on emergency medicine boards questions: an observational study2025
  3. 3Custom GPTs Enhancing Performance and Evidence Compared with GPT-3.5, GPT-4, and GPT-4o? A Study on the Emergency Medicine Specialist Examination2024 · 20 citations
  4. 4Artificial intelligence in medical education2026
  5. 5ChatGPT-4 in comparison with traumatological junior doctors in emergency room cases at a level 1 trauma centre – a pilot study2026