Mixed-methods research evaluates AI scoring reliability for writing assignments, suggesting implications for AI grading systems.
As Artificial Intelligence (AI) is being used more and more in education, utilizing AI to grade the writing of students is a concern for trustworthiness. Using a mixed-methods research design that combines both quantitative and qualitative data collection tools -questionnaires and semi-structured interviews - this study investigates the reliability of using ChatGPT to mark students' writing assignments on the EOP online learning platform (https://eop.edu.vn/) compared with human evaluators at the School of Languages and Tourism, Hanoi University of Industry. The findings provide the advantages and limitations of AI-supported grading, highlighting the accuracy, consistency, and alignment with human grading criteria of AI grading. The attitudes of teachers toward AI scoring are also examined in this paper to determine its accuracy. Recommendations for enhancing AI scoring systems to enable more effective and fairer assessments are provided based on the findings. The research contributes to the academic literature on the use of AI in education, emphasizing the importance of sustaining the enhancement of AI-driven evaluation tools to enable effective and fairer online learning.
No takes yet. Share an insight, caveat, or question.
Tran et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: