Survey explores automatic question generation using large language models in education, highlighting methods and challenges.
Test-based learning is effective in fostering knowledge retention, but manually creating assessment questions remains time-consuming and limits personalized student practice. The advent of Large Language Models (LLMs) has introduced new possibilities for Automatic Question Generation (AQG). Motivated by this context, this survey provides a comprehensive overview of AQG using LLMs, focusing on educational applications. Following the PRISMA methodology, we reviewed 132 studies published between 2023 and 2025. Our contributions include a taxonomy of question types by response openness, an analysis of AQG efforts across knowledge fields, educational levels, evaluation strategies, and difficulty control. We also identify recurring challenges and research opportunities.
No takes yet. Share an insight, caveat, or question.
Oliveira et al. (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: