PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 20250 citationsOpen Access

Assessing Large Language Models for Automated Feedback Generation in Learning Programming Problem Solving

View Full Paper
PSPriscylla SilvaECEvandro Costa

Key Points

  • Accurate feedback generation is crucial for programming education, and LLMs can automate this process.
  • In testing, LLMs provided accurate and complete feedback in 63% of cases, while 37% had mistakes.
  • The study involved four LLM models evaluated against a benchmark dataset of 45 student solutions.
  • Findings call for enhancements to improve the reliability of LLMs in educational applications.

Abstract

Providing effective feedback is important for student learning in programming problem-solving. In this sense, Large Language Models (LLMs) have emerged as potential tools to automate feedback generation. However, their reliability and ability to identify reasoning errors in student code remain not well understood. This study evaluates the performance of four LLMs (GPT-4o, GPT-4o mini, GPT-4-Turbo, and Gemini-1.5-pro) on a benchmark dataset of 45 student solutions. We assessed the models' capacity to provide accurate and insightful feedback, particularly in identifying reasoning mistakes. Our analysis reveals that 63\% of feedback hints were accurate and complete, while 37\% contained mistakes, including incorrect line identification, flawed explanations, or hallucinated issues. These findings highlight the potential and limitations of LLMs in programming education and underscore the need for improvements to enhance reliability and minimize risks in educational applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Silva et al. (2025) studied this question.

synapsesocial.com/papers/68e62de1a8c0c6d45873fe7dhttps://doi.org/10.48550/arxiv.2503.14630
Ask AI
Helpful
Bookmark
Share
View Full Paper