PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 15, 2026Emergency Medicine International0 citationsOpen Access

Accuracy and Reliability of AI Models in Emergency Myocardial Infarction Education

View Full Paper
İKİlhan Korkmaz

Key Points

  • The aim is to evaluate the accuracy, reliability, and readability of AI models in educating about acute myocardial infarction.
  • Cross-sectional study conducted between February and March 2025.
  • Three large language models—ChatGPT-4o, Claude 3.7 Sonnet, and Gemini Advanced 2.0 Flash—were assessed using 30 patient-focused questions.
  • Responses were evaluated by emergency medicine experts on accuracy, reliability (DISCERN, EQIP), and readability indices.
  • ChatGPT-4o achieved the highest accuracy score of 4.38 ± 0.38, significantly outperforming Claude 3.7 (4.09 ± 0.55) and Gemini 2.0 (3.92 ± 0.41) with p < 0.001.
  • ChatGPT-4o showed superior performance in general information and diagnostics with p values of 0.002 and 0.009, respectively.
  • Claude 3.7 produced more readable responses across all indices, significantly better than the other models with p values ≤ 0.003.

Abstract

Background Acute myocardial infarction (AMI) is a major global cause of morbidity and mortality. Large language models (LLMs) are emerging tools for patient education. This study evaluated the performance of three LLMs in delivering accurate, reliable, and readable educational content regarding AMI. Methods In this cross‐sectional study (February–March 2025), a clinical case of a patient with an inferior STEMI ECG was presented to three LLMs: ChatGPT‐4o, Claude 3.7 Sonnet, and Gemini Advanced 2.0 Flash. Each model answered 30 patient‐focused questions across three domains: general disease knowledge, diagnostic processes, and treatment approaches. Responses were assessed by four emergency medicine associate professors (10–20 years of experience) using a 5‐point Likert scale for accuracy, DISCERN and EQIP tools for reliability and quality, and standard readability indices. Results ChatGPT‐4o achieved the highest accuracy score (4.38 ± 0.38), followed by Claude 3.7 (4.09 ± 0.55) and Gemini 2.0 (3.92 ± 0.41) ( p < 0.001). ChatGPT‐4o performed significantly better in general information ( p = 0.002) and diagnostics ( p = 0.009), while Claude 3.7 excelled in treatment‐related content ( p = 0.015). Claude 3.7 produced significantly more readable responses than both ChatGPT‐4o and Gemini 2.0 across all indices (Flesch–Kincaid, p = 0.002, Gunning Fog, p ≤ 0.001, Coleman–Liau, p = 0.003). ChatGPT‐4o scored “excellent” on the DISCERN scale; all models were rated as “good quality with minor shortcomings” on EQIP. Reliability scores did not differ significantly (DISCERN, p = 0.188; EQIP, p = 0.935). Conclusions LLMs show promise in supporting patient education on AMI. While ChatGPT‐4o offers superior accuracy and reliability, Claude 3.7 enhances accessibility through clearer language. This is the first study comparing three LLMs for AMI education in an emergency context, underscoring that physician oversight remains essential for educational applications in emergency medicine.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

İlhan Korkmaz (2026) studied this question.

synapsesocial.com/papers/6a2f96eca1cfeec4908281cfhttps://doi.org/10.1155/emmi/5530861
Ask AI
Helpful
Bookmark
Share
View Full Paper