PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 20250 citationsOpen Access

Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning

View Full Paper
BYBo YuanThe University of QueenslandJHJun HuUniversity of Electronic Science and Technology of China

Key Points

  • GPT-4o produces more informative and structured feedback compared to other models, enhancing personalized guidance.
  • The study involved a realistic tutoring task, utilizing a dataset with student responses to evaluate model performance.
  • Pairwise comparisons conducted with Gemini as a virtual judge revealed differential strengths among the large language models.
  • Results support the potential of large language models as effective educational assistants for individualized learning support.

Abstract

While Large Language Models (LLMs) are increasingly envisioned as intelligent assistants for personalized learning, systematic head-to-head evaluations within authentic learning scenarios remain limited. This study conducts an empirical comparison of three state-of-the-art LLMs on a tutoring task that simulates a realistic learning setting. Using a dataset comprising a student's answers to ten questions of mixed formats with correctness labels, each LLM is required to (i) analyze the quiz to identify underlying knowledge components, (ii) infer the student's mastery profile, and (iii) generate targeted guidance for improvement. To mitigate subjectivity and evaluator bias, we employ Gemini as a virtual judge to perform pairwise comparisons along various dimensions: accuracy, clarity, actionability, and appropriateness. Results analyzed via the Bradley-Terry model indicate that GPT-4o is generally preferred, producing feedback that is more informative and better structured than its counterparts, while DeepSeek-V3 and GLM-4.5 demonstrate intermittent strengths but lower consistency. These findings highlight the feasibility of deploying LLMs as advanced teaching assistants for individualized support and provide methodological guidance for future empirical research on LLM-driven personalized learning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yuan et al. (2025) studied this question.

synapsesocial.com/papers/68ec384042a190b2c351985bhttps://doi.org/10.48550/arxiv.2509.05346
Ask AI
Helpful
Bookmark
Share
View Full Paper