PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 13, 2025Heart & Lung4 citationsOpen Access

Comparing DeepSeek and GPT-4o in ECG interpretation: Is AI improving over time?

SGSerkan GünayAÖAhmet ÖztürkAKAnılcan Tahsin Karahan

Structured PICO

Does DeepSeek provide accurate ECG interpretation compared to GPT-4o and human experts?

P
Population
40 ECG images (20 daily routine, 20 more challenging) from the book 150 ECG Cases
I
Intervention
DeepSeek large language model (evaluated 13 times per ECG)
C
Comparator
GPT-4o (evaluated 13 times per ECG), 12 cardiologists, and 12 emergency medicine specialists
O
Outcome
Accuracy in ECG interpretation (median correct answers)surrogate

GPT-4o outperforms DeepSeek in ECG interpretation accuracy, but both AI models still fall short of expert-level performance.

Abstract

BACKGROUND: DeepSeek is a recently launched large language model (LLM), whereas GPT-4o is an advanced ChatGPT version whose electrocardiography (ECG) interpretation capabilities have been previously studied. However, DeepSeek's performance in this domain remains unexplored. OBJECTIVES: This study aims to evaluate DeepSeek's accuracy in ECG interpretation and compare it with GPT-4o, emergency medicine specialists, and cardiologists. A secondary aim is to assess any performance changes in GPT-4o over one year. METHODS: Between February 9 and March 1, 2025, 40 ECG images (20 daily routine, 20 more challenging) from the book 150 ECG Cases were evaluated by both GPT-4o and DeepSeek, each model tested 13 times. The accuracy of their responses was compared with previously collected answers from 12 cardiologists and 12 emergency medicine specialists. GPT-4o's 2025 performance was compared to its 2024 results on identical ECGs. RESULTS: GPT-4o outperformed DeepSeek with higher median correct answers on daily routine (14 vs. 12), more challenging (13 vs. 10), and total ECGs (27 vs. 22) with statistically significant differences (p=0.048, p<0.001, p<0.001). A moderate agreement was observed between the responses provided by GPT-4o (p<0.001, Fleiss Kappa=0.473), while a substantial agreement was observed in the responses provided by DeepSeek (p<0.001, Fleiss Kappa=0.712). No significant year-over-year improvement was observed in GPT-4o's performance. CONCLUSION: This first evaluation of DeepSeek in ECG interpretation reveals its performance is lower than that of GPT-4o and human experts. While GPT-4o demonstrates greater accuracy, both models fall short of expert-level performance, underscoring the need for caution and further validation before clinical integration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Günay et al. (2025) studied this question.

synapsesocial.com/papers/6a028391bc3ffe278e6506f8https://doi.org/10.1016/j.hrtlng.2025.08.007
Ask AI
Helpful
Bookmark
Share
View Full Paper