PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 27, 2026Sakarya University Journal of Computer and Information Sciences0 citationsOpen Access

Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives

View Full Paper
AKAbdulkadir KaraSYSerkan Yıldırım

Key Points

  • This study aims to assess the effectiveness and fairness of large language models in scoring short answers in Turkish assessments.
  • Evaluated GPT, Gemini, Gemma, and LLaMA models under zero-shot and one-shot conditions with rubric support.
  • Analyzed model performance through internal consistency and decision reliability measures.
  • Conducted error direction analyses to identify tendencies in scoring.
  • LLMs demonstrated high internal consistency but varying decision reliability based on prompt format.
  • Formulated rubrics with clear performance indicators improved model-human alignment and assessment fairness.
  • Models exhibited systematic low-scoring tendencies, indicating potential biases.

Abstract

This study examines the performance of large language models (LLMs) in Turkish short-answer assessments within the measurement and evaluation theory framework. The GPT, Gemini, Gemma, and LLaMA models were evaluated under zero-shot and one-shot conditions with rubric support. The results show that LLMs have high internal consistency, but decision reliability can vary depending on prompt format and example sensitivity. Formulating rubrics with clear and concrete performance indicators increases model-human alignment and assessment fairness. Furthermore, error direction analyses revealed that models can exhibit systematic low-scoring tendencies. The results indicate that LLMs can support teachers in formative assessment with properly structured rubrics, but ethical oversight and pedagogical responsibility remain indispensable in final decisions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kara et al. (2026) studied this question.

synapsesocial.com/papers/6a3f69a5aea7db3c195406echttps://doi.org/10.35377/saucis...1835608
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Large Language Models as Mediators: Addressing Rater Disagreement in Turkish Essay Scoring2025
  2. 2A Framework for Evaluation of Large Language Models in Essay Assessment: Reliability, Alignment, and Causal Reasoning2026 · 1 citations
  3. 3Evaluating large language models for rubric-based essay grading in an undergraduate biology course2026
  4. 4Evaluating Reliability and Bias of Large Language Models in Automated Essay Scoring2026
  5. 5Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability2024 · 128 citations