Evaluation shows fine-tuned AI enhances question generation and grading accuracy in programming education, indicating improved usability.
This study examines the effectiveness of a fine-tuned generative AI system—trained with a domain question bank—for question generation and automated grading in programming education, and evaluates its instructional usability. Methodologically, we constructed an annotated question bank covering nine item types and, under a controlled environment, compared pre- and post-fine-tuning performance on question-type recognition and answer grading using Accuracy, Macro Precision, Macro Recall, and Macro F1. We also collected student questionnaires and open-ended feedback to analyze subjective user experience. Results indicate that the accuracy of question-type recognition improved from 0.6477 to 0.8409, while grading accuracy increased from 0.9474 to 0.9605. Students’ subjective perceptions aligned with these quantitative trends, reporting higher ratings for grading accuracy and question generation quality; overall interactive experience was moderately high, though system speed still requires improvement. These findings provide course-aligned empirical evidence that fine-tuning with domain data can jointly enhance the effectiveness and usability of both automatic question generation and automated grading.
No takes yet. Share an insight, caveat, or question.
Lai et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: