PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 13, 2025Cureus9 citationsOpen Access

Performance of ChatGPT and Large Language Models on Medical Licensing Exams Worldwide: A Systematic Review and Network Meta-Analysis With Meta-Regression

View Full Paper
AKAlousious KasaggaPeking UniversityASAmy R. SapkotaUniversity of Maryland, College ParkGCGichin ChangaramkumarathChongqing Medical University

Key Points

  • Most large language models achieved over 60% accuracy on medical licensing exams, with GPT-o1 leading at 95.4%.
  • The study analyzed 120 evaluations across 10 exam systems and identified significant performance variations by language and model.
  • Meta-regression revealed model type and language significantly impacted exam performance, with lower scores in Chinese and Japanese exams.
  • Heterogeneity among results was largely explained by model version and exam system, highlighting the importance of standardized assessments.

Abstract

Large language models (LLMs) are increasingly being tested on national medical licensing examinations, yet existing research is fragmented across models, exam systems, and languages. This study is the first meta-analysis to systematically assess LLM performance across multiple medical licensing exams and languages using pooled estimates, network meta-analysis, and moderator-aware meta-regression. We synthesized accuracy data from 120 evaluations covering 10 exam systems in nine languages, identified through comprehensive searches of PubMed, Web of Science, and Institute of Electrical and Electronics Engineers (IEEE) Xplore, covering the period from 2021 to June 2025. The random-effects meta-analysis showed that 13 of 16 models exceeded the 60% passing threshold, with GPT-o1 leading at 95.4%, followed by DeepSeek-R1 (92.0%) and GPT-4o (89.4%). P-score rankings and network meta-analysis confirmed the superior performance of GPT-o1, while GPT-3.5 and LLaMA-13B consistently underperformed. Meta-regression revealed a significant variation in accuracy by model version, exam system, and language, with lower performance in Chinese and Japanese exams and higher performance in German and Peruvian settings. After adjustment for exam and language, GPT-o1 and DeepSeek-R1 achieved similar accuracy, both significantly higher than GPT-4. Model type, exam system, and language explained most of the between-study heterogeneity (R² =88.99%), and sensitivity analyses supported the robustness of the pooled estimates. Overall, several LLMs now approach or exceed the accuracy required to pass standardized medical licensing exams, supporting their potential role in medical education and decision support.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kasagga et al. (2025) studied this question.

synapsesocial.com/papers/68ec51df42911f61ef8b201bhttps://doi.org/10.7759/cureus.94300
Ask AI
Helpful
Bookmark
Share
View Full Paper