With the advent of artificial intelligence, large language model (LLM) based Automated Essay Scoring (AES) systems have been developed that can consistently make human-like decisions that do not depend fully on surface level linguistic features. However, research into the use of LLM-based AES systems is limited and little is known about the reliability, agreement, or validity of the systems. The goal of this study was to provide evidence for the reliability, agreement, and validity of LLM-based AES systems in a standardized writing assessment used for secondary school students. Both representation and generative LLM-based AES systems were developed to score persuasive essays and assessed for reliability. Then the agreement of the developed AES systems with human raters was assessed through correlational analyses. We used extrinsic convergent validation approaches to examine if the human and LLM scores correlated with linguistic components. Results indicate strong reliability and agreement for the LLM scores. In terms of convergent validity, initial correlational analyses indicated that the representation LLM AES system showed differential correlations with the human scores in terms of a text length and type-token ratio component. This result contrasts with the correlational results from the generative LLM AES model, which indicated no differences in associations between the model and human scores with regards to the linguistic components.
CROSSLEY et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: