PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 8, 20260 citationsOpen Access

Comparing Zero-Shot Large Language Model Prompting with Human Coding of Theory Concepts in Student Essays

View Full Paper
SKShelley KeithPPPhilip I. PavlikKSKristen L. Stives

Key Points

  • This research aims to understand how large language models perform in coding student essays compared to human coders.
  • Compared coding accuracy for four LLMs and human coding.
  • Analyzed human-AI correlations and biases across different models and prompts.
  • Examined performance across various criminological theories and content dimensions.
  • LLM choice significantly affected human-AI correspondence, with Claude Sonnet 4 performing best.
  • Prompt type did not greatly influence overall performance.
  • Error rates were lowest for identifying listed concepts and highest for definition accuracy.

Abstract

Recent studies have explored the cost and time benefits of using artificial intelligence (AI), particularly large language models (LLMs), in coding student essays. While these models show promise, not enough is understood about the factors that affect how their qualitative coding performance compares to human coding. This study examines coding accuracy for content errors in college student essays on criminological theories by comparing human-coded results with outputs from four LLMs. We evaluated human-AI correlations, AI error, and AI bias across four LLMs, five prompt types, three theory content coding dimensions, and four criminological theories. Results indicate that LLM choice significantly influenced human-AI correspondence, with Claude Sonnet 4 exhibiting the best overall performance and GPT 4.1 Mini the worst. Prompt type had minimal impact on performance. Across models, error rates were lowest when identifying whether students listed a concept, and highest when assessing whether definitions were correct. LLMs performed better on concise theories than on more complex ones. The code is available at https://github.com/imrryr/LLM-queries

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Keith et al. (2026) studied this question.

synapsesocial.com/papers/69d5f13674eaea4b11a7abd7https://doi.org/10.5281/zenodo.19443161
Ask AI
Helpful
Bookmark
Share
View Full Paper