Background: Large language models (LLMs) are increasingly being explored in medical education and may support clinical knowledge acquisition. However, evidence regarding their performance in non-English, specialty-level national medical examinations remains limited. Objective: This study aimed to compare the performance of five LLMs in the Polish State Specialization Examination (Państwowy Egzamin Specjalizacyjny (PES)) in anesthesiology and intensive care medicine and to assess the association between empirical question difficulty and model accuracy. Methods: We evaluated ChatGPT 5.2 (OpenAI, San Francisco, CA, USA), Grok 3 (xAI, San Francisco, CA, USA), Claude Sonnet 4.5 (Anthropic, San Francisco, CA, USA), DeepSeek R1 (DeepSeek, Hangzhou, China), and Gemini 2.5 Pro (Google DeepMind, Mountain View, CA, USA) using questions from the autumn 2024 and spring 2025 sessions of the PES in anesthesiology and intensive care medicine. Each session consisted of 120 single-best-answer multiple-choice questions in Polish. The models were prompted to select only one answer option, from A to E. Model accuracy was compared using Cochran’s Q test, followed by post hoc McNemar tests with Holm correction. The association between the difficulty index, derived from examinee performance, and model accuracy was assessed using Spearman’s rank correlation and odds ratios (ORs). Results: In the autumn 2024 session, model accuracy ranged from 79.2% to 85.8%, with no statistically significant global differences between models. Gemini 2.5 Pro achieved the highest score, answering 103 of 120 questions correctly. In the spring 2025 session, accuracy ranged from 81.7% to 93.3%, and significant global differences were observed between models. Gemini 2.5 Pro achieved the best result, with 112 correct answers out of 120, and significantly outperformed ChatGPT 5.2 and Claude Sonnet 4.5 after the Holm correction. Higher values of the difficulty index were associated with greater model accuracy in both examination sessions. Model performance was higher for questions with difficulty index ≥ 0.5 than for questions with difficulty index < 0.5, 87.4% versus 66.3%, respectively (OR 3.53; 95% CI 2.50-4.99; Wald z = 7.16; p = 7.71 × 10⁻¹³). Conclusions: Contemporary LLMs achieved high accuracy on Polish specialty-level examination questions in anesthesiology and intensive care medicine. These findings suggest potential educational utility in specialty examination preparation. However, high performance on multiple-choice questions should not be interpreted as evidence of readiness for autonomous clinical use.
Wojtas et al. (Fri,) studied this question.