Objective Structured Clinical Examinations (OSCEs) are central to assessing medical student clinical competence, but human grading imposes substantial burden. We report, to our knowledge, the first prospective single-center deployment of an integrated multimodal artificial intelligence (AI) system for OSCE grading in undergraduate medical education, spanning notes, audio, and video. We present MAPLES (Multimodal Assessment Pipeline for Learning Encounter Scoring), a rubric-driven zero-shot multimodal LLM system deployed in Fall 2025 for 222 students (72,907 retained item-level scores). In the routed low-scoring review set, tolerant AI–standardized patient evaluator (SPE) agreement was 85.7–92.4% by modality. In a selected set of 616 disagreements adjudicated with visible score provenance, physician scores matched AI on 76.0% of items and the SPE on 19.0%. Human scoring passes fell by 92.3% versus a modeled single-pass manual comparator. Together, these findings show that multimodal AI first-pass scoring can be embedded in routine OSCE operations, with human review on the low-scoring tail and physician adjudication of selected disagreements.
No takes yet. Share an insight, caveat, or question.
Ngo et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: