PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 16, 2025Applied Sciences11 citationsOpen Access

A Comparative Study of Large Language Models in Programming Education: Accuracy, Efficiency, and Feedback in Student Assignment Grading

View Full Paper
ABAndrija BernikDRDanijel RadoševićAČAndrej Čep

Key Points

  • AI-assisted grading improves evaluation speed while ensuring high-quality feedback for students.
  • Quantitative analysis shows high correlation with instructor evaluations, with ChatGPT-4 scoring highest with a Pearson coefficient of 0.91.
  • Utilizing structured prompts aligned with rubrics enhances the consistency and reliability of AI evaluations.
  • Challenges include addressing model biases, ensuring academic integrity, and emphasizing the need for human oversight.

Abstract

Programming education traditionally requires extensive manual assessment of student assignments, which is both time-consuming and resource-intensive for instructors. Recent advances in large language models (LLMs) open opportunities for automating this process and providing timely feedback. This paper investigates the application of artificial intelligence (AI) tools for preliminary assessment of undergraduate programming assignments. A multi-phase experimental study was conducted across three computer science courses: Introduction to Programming, Programming 2, and Advanced Programming Concepts. A total of 315 Python assignments were collected from the Moodle learning management system, with 100 randomly selected submissions analyzed in detail. AI evaluation was performed using ChatGPT-4 (GPT-4-turbo), Claude 3, and Gemini 1.5 Pro models, employing structured prompts aligned with a predefined rubric that assessed functionality, code structure, documentation, and efficiency. Quantitative results demonstrate high correlation between AI-generated scores and instructor evaluations, with ChatGPT-4 achieving the highest consistency (Pearson coefficient 0.91) and the lowest average absolute deviation (0.68 points). Qualitative analysis highlights AI’s ability to provide structured, actionable feedback, though variability across models was observed. The study identifies benefits such as faster evaluation and enhanced feedback quality, alongside challenges including model limitations, potential biases, and the need for human oversight. Recommendations emphasize hybrid evaluation approaches combining AI automation with instructor supervision, ethical guidelines, and integration of AI tools into learning management systems. The findings indicate that AI-assisted grading can improve efficiency and pedagogical outcomes while maintaining academic integrity.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bernik et al. (2025) studied this question.

synapsesocial.com/papers/68d4508931b076d99fa585b3https://doi.org/10.3390/app151810055
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The future of grading programming assignments in education: The role of ChatGPT in automating the assessment and feedback process2024 · 123 citations
  2. 2Time to Revisit Existing Student’s Performance Evaluation Approach in Higher Education Sector in a New Era of ChatGPT — A Case Study2023 · 252 citations
  3. 3Feedback sources in essay writing: peer-generated or AI-generated feedback?2024 · 263 citations
  4. 4Reliability of ChatGPT in automated essay scoring for dental undergraduate examinations2024 · 71 citations
  5. 5LLM-based automatic short answer grading in undergraduate medical education2024 · 86 citations