Key points are not available for this paper at this time.
This study examines the accuracy and fairness of generative AI–based automated essay scoring (AES) in a developmentally emergent writing context, compared with a feature-based scoring engine and human ratings. Despite growing interest in generative AI for AES, limited research has examined model performance in early-grade writing, where linguistic and transcription skills are still developing. Using 1768 essays from Grades 3–4, we evaluated GPT-4o under three prompting strategies and a fine-tuned configuration. Human–machine agreement was assessed against human–human benchmarks at both holistic and trait levels using quadratic weighted kappa, exact and adjacent agreement rates, and classification metrics. Fairness was examined using overall score accuracy, overall score difference, and conditional score difference. Results indicate that carefully designed prompting improves GPT-4o’s scoring accuracy, approaching human and feature-based AES performance, while the fine-tuned model achieved the closest alignment to human–human agreement. Trait-level analyses revealed complementary strengths: the best-performing prompting approach aligned more closely on higher-level traits (e.g., organization and style), whereas the feature-based engine showed stronger alignment on development of ideas and all surface-level traits. Fairness analyses indicated that no model was entirely free of subgroup differences. These findings suggest that generative AI–based AES can approximate human scoring in early-grade writing when appropriately calibrated.
Yue et al. (Fri,) studied this question.