The rapid growth of programming education and online learning environments has increased the demand for scalable and reliable assessment methods. Although many automated grading approaches exist, the literature spans traditional test-based methods and recent solutions based on artificial intelligence, making it difficult to obtain a coherent overview of the field. This study conducts a systematic literature review on automated programming assessment, analysing grading approaches, the use of artificial intelligence techniques, educational contexts, and comparisons between automated and human grading. The review synthesises evidence from 39 primary studies to identify technological trends and research gaps. The results show a shift from execution-based grading toward deep learning and large language model approaches, with most systems applied in higher education and limited research on visual programming assessment. Automated grading is evolving into an intelligent educational support tool, but challenges remain with regard to reliability and pedagogical impact. Future work should prioritise empirical validation, hybrid grading models, and broader educational applications.
Mušac et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: