PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 20260 citationsOpen Access

Using LLMs to Identify Indicators of Persistence from Students' Dialogues with a Pedagogical Agent

View Full Paper
TOTeresa OberSZShan ZhangCFCarol Forsyth

Key Points

  • The aim is to utilize large language models to code indicators of persistence and related constructs in student interactions.
  • Analyzed chat log data from 107 middle school students during a conversation-based assessment in mathematics.
  • Implemented Large Language Models with various configurations of temperature settings and model types for analysis.
  • Evaluated LLM performance against human expert coders to assess reliability and alignment.
  • Self-efficacy exhibited the strongest alignment between human coders and LLM outputs.
  • Higher temperatures improved coding accuracy for constructs with moderate coherence.
  • Deterministic settings were more effective for well-defined constructs.

Abstract

Conversational learning systems offer new opportunities to examine learning processes through chat log data. Constructs such as persistence, self-efficacy, interest, perceived challenge, and prior knowledge are known predictors of student performance but are challenging to detect at scale using traditional methods. This study explores the use of Large Language Models (LLMs) to automatically code indicators of these constructs from student chat logs collected through a conversation-based assessment (CBA) for middle school mathematics. Indicators included observable behaviors such as students' expressions of challenge, help-seeking, goal-setting, and self-regulatory strategies evident in their conversational interactions within the CBA. We evaluated multiple configurations of ChatGPT4o, varying temperature settings (0, .3, .7, 1) and model types (mini vs. regular), against human expert coders. The dataset comprised over 10,000 student turns collected from 107 middle school students classified as English learners as they interact with the CBA. Reliability was assessed within and between LLM configurations and humans. Results reveal systematic patterns: constructs with moderate theoretical coherence benefited from higher temperatures, while well-defined constructs required deterministic settings. Self-efficacy showed the highest human-LLM alignment. These findings illustrate the challenges of measuring complex psychological constructs and highlight the promise of human-LLM collaboration to enhance qualitative coding efficiency and validity in educational research. Supplemental materials are available online here: https://doi.org/10.17605/osf.io/s85ck.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ober et al. (2026) studied this question.

synapsesocial.com/papers/69a91e4cd6127c7a504c2301https://doi.org/10.5281/zenodo.18852440
Ask AI
Helpful
Bookmark
Share
View Full Paper