PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026Educational Measurement Issues and Practice1 citationsOpen Access

AI‐Generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity

View Full Paper
YZYang ZhongJHJiangang HaoMFMichael Fauss

Key Points

  • The central aim is to analyze the characteristics of AI-generated essays and their effects on assessment processes.
  • Analyzed large-scale empirical data involving essays from various LLMs.
  • Benchmarking the quality and characteristics of AI-generated essays.
  • Evaluated existing automated scoring systems and their limitations.
  • Identified potential improvements for scoring features and weight recalibration.
  • Current automated scoring systems face limitations with AI-generated essays.
  • Detectors trained on essays from one model can accurately identify texts from other models.
  • Found opportunities for improving detection methods to enhance academic integrity.

Abstract

Abstract The rapid advancement of large language models (LLMs) has enabled the generation of coherent essays, making AI‐assisted writing increasingly common in educational and professional settings. Using large‐scale empirical data, we examine and benchmark the characteristics and quality of essays generated by popular LLMs and discuss their implications for two key components of writing assessments: automated scoring and academic integrity. Our findings highlight limitations in existing automated scoring systems when applied to essays generated or heavily influenced by AI, and identify areas for improvement, including the development of new features to capture deeper thinking and recalibrating feature weights. Despite growing concerns that the increasing variety of LLMs may undermine the feasibility of detecting AI‐generated essays, our results show that detectors trained on essays generated from one model can often identify texts from others with high accuracy, suggesting that effective detection could remain manageable in practice.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhong et al. (2026) studied this question.

synapsesocial.com/papers/6980ff08c1c9540dea811b0ehttps://doi.org/10.1111/emip.70013
Ask AI
Helpful
Bookmark
Share
View Full Paper