Key points are not available for this paper at this time.
BACKGROUND: Generative artificial intelligence (GenAI) technologies —exemplified by large language models (LLMs) such as ChatGPT and DeepSeek—have attracted substantial interest in medical education for their capacity to generate content, adapt to individual learners, and support interactive dialogue. Whether these technologies produce measurable improvements in educational outcomes remains uncertain, and this evidence gap has constrained broader adoption. We conducted this systematic review and meta-analysis of randomized controlled trials (RCTs) to evaluate the effects of GenAI-assisted teaching on knowledge acquisition, clinical skills, and learner attitudes. METHODS: We conducted a systematic review and meta-analysis of randomized controlled trials (RCTs) following the PRISMA 2020 guidelines. The study protocol was registered on PROSPERO (CRD420261277187).Six databases (PubMed, Cochrane Library, Embase, Web of Science, CNKI, and Wanfang) were searched from January 2020 to February 2026. We included RCTs in which GenAI-integrated medical education was compared against traditional teaching or non-GenAI digital teaching, with participants being medical students, interns, or residents. Risk of bias was evaluated with the Cochrane RoB 2 tool; the GRADE framework was used to judge evidence certainty. Pooling was done through random-effects models with REML estimation. RESULTS: We identified 20 eligible RCTs (1413 participants in total). Post-intervention knowledge scores were higher in GenAI groups (SMD = 0.99, 95% CI: 0.49–1.49, P = 0.001; k = 11; I2 = 86.7%; GRADE: low certainty), and so were clinical skill scores (SMD = 1.09, 95% CI: 0.81–1.38, P < 0.001; k = 8; I2 = 41.0%; GRADE: moderate certainty). At follow-up, knowledge retention remained significant (SMD = 0.51, 95% CI: 0.24–0.78, P = 0.010; k = 4; I2 = 0%; GRADE: moderate certainty). Publication bias was not detected by Egger’s test (P = 0.104); trim-and-fill adjustment gave a similar pooled estimate (SMD = 0.90). Subgroup comparisons by AI tool type, learner level, interaction modality, and risk of bias were all non-significant. CONCLUSION: GenAI-assisted teaching was associated with moderate-to-large pooled standardised mean differences for clinical skills (SMD = 1.09; moderate certainty) and post-intervention knowledge (SMD = 0.99; low certainty). However, the prediction interval for knowledge crossed zero (− 0.57 to 2.56), indicating that the direction of benefit cannot be assumed in future contexts. In four studies, knowledge gains were observed at follow-up (SMD = 0.51; moderate certainty), though the small number of studies limits confidence in durability. Evidence certainty is low to moderate, owing to high heterogeneity, risk of bias, and the absence of blinding. Current evidence is insufficient to draw definitive conclusions; GenAI may serve as a useful supplement to conventional instruction but should not be adopted as a replacement without evidence from well-powered, pre-registered, multi-centre trials. TRIAL REGISTRATION: PROSPERO registration: CRD420261277187.
Wang et al. (Thu,) studied this question.