Although large language models (LLMs) show potential in medical education, their effectiveness in Chinese anesthesiology Standardized Residency Training Program (SRTP) exams remains unexplored. This study aimed to assess the performance, consistency, and clinical reasoning capabilities of LLMs in this specific context. We conducted a multidimensional evaluation of three LLMs (GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro) using a 210-question Chinese mock SRTP exam. The exam encompassed four types of questions with increasing complexity: A1 (knowledge-recall), A2 (knowledge-application), A3/A4 (multi-step clinical scenarios), and case-analysis (complex reasoning with partial-credit scoring). Each model was tested across 30 repeated iterations to evaluate both accuracy and consistency. Their performance was compared to anonymized results from 32 human SRTP trainees. All LLMs exceeded the passing threshold (400/650 points). Among them, Claude 3.5 Sonnet achieved the highest mean score (495.4 ± 9.5), followed by Gemini 1.5 Pro (493.8 ± 4.3) and GPT-4o (482.4 ± 10.9). Gemini 1.5 Pro outperformed the others on A1 questions (82.5% median accuracy, P 0.05 vs. Gemini 1.5 Pro) and 42.7% for GPT-4o ( P = 0.04 vs. Gemini 1.5 Pro). Notably, all LLMs outperformed human trainees, who had an average score of 426.8 ± 25.9 (all P values < 0.001), although the gap narrowed for the most complex questions. State-of-the-art LLMs demonstrate high proficiency on Chinese anesthesiology SRTP exams, surpassing human performance in structured assessments. Their strengths in knowledge recall and potential for scalable feedback make them promising adjuncts for mitigating training disparities. However, the diminished performance in complex clinical reasoning tasks suggests that these models should complement rather than replace human-centric education.
Wang et al. (Sat,) studied this question.