PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 2, 2026Computers and Education Open0 citationsOpen Access

Comparing GPT and human raters in essay assessment: Variability, bias, and the potential of LLM-based scoring

View Full Paper
HWHsiang-Ning WuMCMan-ni ChuJHJia-Lien Hsu

Key Points

  • The study aims to evaluate the performance of a GPT-based scoring system alongside human raters in assessing second-language essays.
  • Analyzed ratings from 20 human raters (10 native and 10 non-native English speakers) and GPT across four criteria.
  • Utilized a Many-Facet Rasch Measurement model to assess variability and bias in scoring.
  • Evaluated 181 university L2 essays based on Content, Organization, Grammar, and Vocabulary.
  • Rater variability and potential bias were observed among human raters.
  • GPT showed high consistency in scoring but had range compression at higher scores compared to human raters.
  • Findings suggest that GPT could complement traditional scoring methods but may miss subjective writing features.

Abstract

Ensuring reliable essay scoring is challenging when rater severity and bias vary by linguistic background. This study investigates whether a GPT-based Automated Essay Scoring (AES) system can complement human raters in second-language (L2) writing assessment. Using a Many-Facet Rasch Measurement (MFRM) model, this study analyzed ratings from 20 human raters—10 native English-speaking (NES) and 10 non-native English-speaking (NNES)—and GPT across four analytic criteria ( Content , Organization , Grammar , and Vocabulary ) on 181 university L2 essays. The results indicate that raters demonstrated variability and potential rater bias; GPT exhibited high internal consistency but range compression at the upper end compared to humans. These findings offer insights into how GPT can supplement or complement traditional rater-based methods, potentially alleviating the time-consuming and subjective aspects of human scoring. Nonetheless, concerns remain regarding its sensitivity to subjective writing features such as creativity, tone, and nuanced lexical use. This study contributes to the growing understanding of large language models in educational contexts and highlights the need for further refinement and validation of AES systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2026) studied this question.

synapsesocial.com/papers/69a52e04f1e85e5c73bf14c4https://doi.org/10.1016/j.caeo.2026.100341
Ask AI
Helpful
Bookmark
Share
View Full Paper