PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 2026Jurnal Pendidikan Teori Penelitian dan Pengembangan0 citationsOpen Access

Reliability of ChatGPT-Based Essay Scoring: A Teacher–AI Comparison in Economics Education

View Full Paper
RPRamadzan Defitri PratamaSebelas Maret UniversityKSKhresna Bayu SangkaSebelas Maret UniversityCICicilia Dyah Sulistyaningrum IndrawatiSebelas Maret University

Key Points

  • The study aims to evaluate how reliable the ChatGPT-based essay scoring system is compared to teacher assessments in economics education.
  • Quantitative approach with a comparative reliability study design
  • Involvement of 60 high school students in grades 10 and 11
  • Data analysis using Intraclass Correlation Coefficient (ICC), Pearson's correlation, and paired sample t-test
  • EsyGrade shows a very high level of reliability in essay scoring
  • Agreement is strong between EsyGrade and teacher assessments
  • Consistency observed in both total scores and individual assessment dimensions

Abstract

Essay assessment in economics education plays an important role in measuring students' conceptual understanding and higher-order thinking skills, but is often hampered by subjectivity and teachers' workload. This study aims to analyze the reliability and level of agreement between teacher assessments and the ChatGPT-based automatic essay grading system (EsyGrade). The study uses a quantitative approach with a comparative reliability study design involving 60 high school students in grades 10 and 11. Data were analyzed using the Intraclass Correlation Coefficient (ICC), Pearson's correlation, and paired sample t-test. The results show that EsyGrade has a very high level of reliability and is consistent with teacher assessments, both in terms of total scores and each dimension of essay assessment. These findings indicate that EsyGrade has the potential to be a reliable and objective essay assessment support tool in economics learning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pratama et al. (2026) studied this question.

synapsesocial.com/papers/69b2579096eeacc4fcec6572https://doi.org/10.17977/2502-471x.1183
Ask AI
Helpful
Bookmark
Share
View Full Paper