PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 17, 2025Computers9 citationsOpen Access

CourseEvalAI: Rubric-Guided Framework for Transparent and Consistent Evaluation of Large Language Models

View Full Paper
CACătălin AnghelMCMarian CrăciunEPEmilia Pecheanu

Key Points

  • The fine-tuned large language model significantly outperformed its base version in rubric-guided evaluations, leading to improved scores.
  • Inter-rater reliability metrics indicated moderate human agreement but less alignment with automated evaluations, suggesting room for improvement.
  • Analyses showed decreased scoring variance and greater clarity in evaluator feedback, enhancing pedagogical alignment in outputs.
  • CourseEvalAI's framework offers a structured methodology for rubric-guided evaluations, promising greater transparency in educational contexts.

Abstract

Background and objectives: Large language models (LLMs) show promise in automating open-ended evaluation tasks, yet their reliability in rubric-based assessment remains uncertain. Variability in scoring, feedback, and rubric adherence raises concerns about transparency and pedagogical validity in educational contexts. This study introduces CourseEvalAI, a framework designed to enhance consistency and fidelity in rubric-guided evaluation by fine-tuning a general-purpose LLM with authentic university-level instructional content. Methods: The framework employs supervised fine-tuning with Low-Rank Adaptation (LoRA) on rubric-annotated answers and explanations drawn from undergraduate computer science exams. Responses generated by both the base and fine-tuned models were independently evaluated by two human raters and two LLM judges, applying dual-layer rubrics for answers (technical or argumentative) and explanations. Inter-rater reliability was reported as intraclass correlation coefficient (ICC(2,1)), Krippendorff’s α, and quadratic-weighted Cohen’s κ (QWK), and statistical analyses included Welch’s t tests with Holm–Bonferroni correction, Hedges’ g with bootstrap confidence intervals, and Levene’s tests. All responses, scores, feedback, and metadata were stored in a Neo4j graph database for structured exploration. Results: The fine-tuned model consistently outperformed the base version across all rubric dimensions, achieving higher scores for both answers and explanations. After multiple-testing correction, only the Generative Pre-trained Transformer (GPT-4)—judged Technical Answer contrast remains statistically significant; other contrasts show positive trends without passing the adjusted threshold, and no additional significance is claimed for explanation-level results. Variance in scoring decreased, inter-model agreement increased, and evaluator feedback for fine-tuned outputs contained fewer vague or critical remarks, indicating stronger rubric alignment and greater pedagogical coherence. Inter-rater reliability analyses indicated moderate human–human agreement and weaker alignment of LLM judges to the human mean. Originality: CourseEvalAI integrates rubric-guided fine-tuning, dual-layer evaluation, and graph-based storage into a unified framework. This combination provides a replicable and interpretable methodology that enhances the consistency, transparency, and pedagogical value of LLM-based evaluators in higher education and beyond.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Anghel et al. (2025) studied this question.

synapsesocial.com/papers/68f199bfde32064e504dcac8https://doi.org/10.3390/computers14100431
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Diagnosing Bias and Instability in LLM Evaluation: A Scalable Pairwise Meta-Evaluator2025 · 13 citations
  2. 2Multi-Model Dialectical Evaluation of LLM Reasoning Chains: A Structured Framework with Dual Scoring Agents2025 · 10 citations
  3. 3Optimizing Large Language Models on Multi-Core CPUs: A Case Study of the BERT Model2024 · 8 citations
  4. 4Medical LLMs: Fine-Tuning vs. Retrieval-Augmented Generation2025 · 27 citations
  5. 5Is GPT-4 a reliable rater? Evaluating consistency in GPT-4's text ratings2023 · 79 citations