PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 20, 2026Discover Artificial Intelligence2 citationsOpen Access

Exploring potential of large language models for automated essay scoring in education

NMNimra MughalAIAli Shariq ImranSDSher Muhammad Daudpota

Key Points

  • The research aims to explore the effectiveness of large language models in automated essay scoring systems.
  • Evaluated LLMs such as GPT and Gemini on benchmark datasets ASAP–AES and LA–AES.
  • Collected real-world data from O-Level classrooms to assess scoring objectivity.
  • Engaged multiple human evaluators to identify biases in traditional grading.
  • Gemini achieved a higher average QWK score of 0.45 on ASAP–AES compared to 0.43 for GPT.
  • The LLM-based scoring showed improved objectivity and reduced bias versus human assessment.

Abstract

Abstract The assessment of open-ended written work is of vital importance to the student learning experience. Conventional essay grading methods heavily depend on expert manual assessment, making them susceptible to errors due to fatigue, bias, and subjectivity. To address this, recent research has introduced AI-based Automated Essay Scoring (AES) systems. While most studies have concentrated on predicting scores, only a few have integrated AES systems with the well-known Large Language Models (LLMs). This study explores the application of LLMs, including GPT and Gemini for AES. The proposed approach was evaluated on two benchmark datasets, namely “Hewlett Foundation: Automated Essay Scoring (ASAP–AES)” and “Learning Agency Lab–Automated Essay Scoring 2.0 (LA–AES)”. The proposed method achieved promising results in AES, demonstrating effectiveness on both the benchmark datasets. Statistical analysis revealed that Gemini outperformed GPT, achieving an average Quadratic Weighted Kappa (QWK) score of 0.45 on the ASAP–AES and 0.43 on the LA–AES. To assess the generalizability and objectivity of the proposed approach, real-world data was collected from an O-Level classroom at Sukkur IBA Community College, Pakistan. Multiple human evaluators participated in the study to examine potential biases in human assessment. The findings indicate that LLM-based scoring demonstrates improved objectivity and reduced bias compared to human assessors.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mughal et al. (2026) studied this question.

synapsesocial.com/papers/6997fa80ad1d9b11b3453cb8https://doi.org/10.1007/s44163-026-01002-y
Ask AI
Helpful
Bookmark
Share
View Full Paper