PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 3, 2025Scientific Reports8 citationsOpen Access

Comparative performance of ChatGPT-4o, ChatGPT-5, and gemini 2.5 flash on Persian internal medicine subspecialty board exams

View Full Paper
SSShahab SheikhalishahiAHAlireza HaddadiSSSaina Sadeghipour

Key Points

  • Gemini 2.5 Flash achieved the highest accuracy of 79.9% in Persian internal medicine subspecialty board exams, highlighting its effectiveness.
  • ChatGPT-5 outperformed ChatGPT-4o with a significant accuracy increase of 74.5% compared to 68.9%, confirming improvements in model development.
  • An artificial neural network combining capabilities of all models reached 81.6% accuracy, suggesting that integrating models can enhance performance.
  • Results emphasize the potential role of AI in medical education and clinical practice but call for further research in practical applications.

Abstract

This study compared the performance of ChatGPT-4o, ChatGPT-5, and Gemini 2.5 Flash on the 2025 Iranian internal medicine subspecialty board examinations. A total of 650 multiple-choice questions from six subspecialties were tested, excluding image-based items. Each question was presented in Persian, and responses were evaluated against the official answer key. Accuracy rates were 68.9% for ChatGPT-4o, 74.5% for ChatGPT-5, and 79.9% for Gemini 2.5 Flash, with Gemini performing significantly better than both ChatGPT versions. ChatGPT-5 also showed a significant improvement over ChatGPT-4o, confirming rapid progress in model development. Subspecialty analysis revealed stronger results in rheumatology and respiratory medicine compared to nephrology, while question type and length had no significant impact on outcomes. An artificial neural network that combined the outputs of all three models reached 81.6% accuracy, slightly exceeding Gemini alone. These findings highlight Gemini-2.5 as the most reliable model for this high-stakes internal medicine exam. The results support the growing role of advanced AI systems as assistants in medical education and clinical practice. However, further research is needed to assess their use in multimodal and real-world clinical tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sheikhalishahi et al. (2025) studied this question.

synapsesocial.com/papers/694025912d562116f28fe93ehttps://doi.org/10.1038/s41598-025-31251-3
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluating the image recognition capabilities of GPT-4V and Gemini Pro in the Japanese national dental examination2024 · 18 citations
  2. 2Evaluating the Performance of State-of-the-Art Artificial Intelligence Chatbots Based on the WHO Global Guidelines for the Prevention of Surgical Site Infection: Cross-Sectional Study2025 · 11 citations
  3. 3Evidence-based potential of generative artificial intelligence large language models in orthodontics: a comparative study of ChatGPT, Google Bard, and Microsoft Bing2024 · 115 citations
  4. 4Comparative performance of ChatGPT, Gemini, and final-year emergency medicine clerkship students in answering multiple-choice questions: implications for the use of AI in medical education2025 · 11 citations
  5. 5Comparative evaluation of AI platforms “Google Gemini 2.5 Flash, Google Gemini 2.0 Flash, DeepSeek V3 and ChatGPT 4o” in solving multiple-choice questions from different subtopics of anatomy2025 · 14 citations