Vision-language models (VLMs) integrate visual and textual information and are increasingly being used as innovative tools in educational applications. However, there is a lack of evidence regarding current practices for integrating VLMs into teaching and learning. To address this research gap and identify the opportunities and challenges associated with the integration of VLMs in education, this paper presents a systematic review of VLM use in formal educational contexts. Peer-reviewed articles published between 2020 and 2025 were retrieved from five major databases: ACM Digital Library, Scopus, Web of Science, Engineering Village, and IEEE Xplore. Following the PRISMA-guided framework, 42 articles were selected for inclusion. Data were extracted and analyzed against six research questions: (1) where VLMs are applied across academic disciplines and educational levels; (2) what types of VLM solutions are deployed and which image–text modalities they infer and generate; (3) the pedagogical roles of VLMs within teaching workflows; (4) reported outcomes and benefits for learners and instructors; (5) challenges and risks identified in practice, together with corresponding mitigation strategies; and (6) reported evaluation methods. The included studies span K-12 through higher education and cover diverse disciplines, with deployments dominated by pre-trained models and a smaller number of domain-adapted approaches. VLM-supported pedagogical functions cluster into five roles: analyst, assessor, content curator, simulator, and tutor. This review concludes by discussing implications for VLM adoption in educational settings and offering recommendations for future research.
Jing Tian (Wed,) studied this question.