PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
Synapse
⌘+K
Synapse
January 24, 2026PLoS ONEOpen Access

Performance of DeepSeek and ChatGPT on the Chinese Health Professional and Technical Examination: A comparative study

View Full Paper
Ask AI
Bookmark
Share

Authors

XLXu LiXHXu HuHXHuiting Xu

Discussion

Loading...

Member takes

Overview

Comparative analysis shows DeepSeek-R1 outperforms GPT-4o API in nursing examination accuracy, indicating model reliability issues.

Key Points

  • To evaluate and compare the performance of DeepSeek-R1 and GPT-4o API on the Chinese Health Professional and Technical Examination.
  • Utilized 400 official multiple-choice practice questions categorized into competency units and types.
  • Assessed overall accuracy, response consistency, and consistent accuracy between the models.
  • Conducted stratified analyses and statistical comparisons using chi-square tests with multiple-comparison corrections.
  • DeepSeek-R1 achieved 88.5% accuracy compared to GPT-4o API's 67.9% (P < 0.001).
  • GPT-4o API showed 96.5% response consistency, while DeepSeek-R1 had 88.5%.
  • DeepSeek-R1 had higher consistent accuracy (84.0%) versus GPT-4o API's 66.7% across several nursing domains.

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/69746090bb9d90c67120a778https://doi.org/10.1371/journal.pone.0338328
View Full Paper
Ask AI
Bookmark
Share