Key points are not available for this paper at this time.
The rapid adoption of generative artificial intelligence (GenAI) in educational institutions has occurred without systematic assessment of its environmental impacts. Training and operating large language models result in electricity consumption, greenhouse gas emissions, water use, hardware manufacturing, and electronic waste, yet these physical footprints remain largely invisible to educators and learners. This PRISMA-based systematic review addressed two research questions: (a) What types of environmental impacts are associated with GenAI tools in educational contexts? and (b) What indicators, data sources, assumptions, and methodological approaches are used to measure or estimate these impacts? A search of four databases (Scopus, Web of Science, ERIC, IEEE Xplore) identified 23 eligible publications from 2022 to 2026, classified into two evidence streams: education-specific studies (n = 8) examined GenAI in direct educational settings, while complementary technical studies (n = 15) provided transferable environmental indicators, benchmarks, and assessment methods from the broader AI sustainability literature. This two-stream design was chosen because education-specific environmental evidence remains scarce, and methodological approaches from technical literature are essential for understanding how educational impacts could be assessed. The synthesis found that operational electricity consumption, carbon emissions, and water demand are the most frequently reported impacts, but evidence in educational settings remains limited and methodologically inconsistent. Technical studies provide precise hardware-level measurements and lifecycle assessments, while educational research relies mainly on indirect proxies, self-reported surveys, or token-count approximations. Education-specific empirical evidence remains limited, with a small number of studies suggesting that GenAI use can increase the energy footprint of student work under some conditions, while real-time carbon-feedback displays may reduce prompt volume and estimated emissions. Comparative labor studies provide mixed findings: although some estimates portray AI-generated work as less carbon-intensive than human labor, correctness-controlled evidence from programming tasks shows that emissions can increase substantially when iterative prompting and output verification are included. Water footprints, embodied emissions, and electronic waste are acknowledged conceptually but measured empirically only in related technical fields. The review proposes a three-tiered policy framework spanning pedagogical practices, technical platform choices, and institutional governance to align educational AI use with environmental sustainability commitments.
Radovan et al. (Wed,) studied this question.