PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026ZDM1 citationsOpen Access

Automated coding of content and pedagogical content knowledge of mathematics using a multi-agent large language model

YCYasemin Copur-GencturkKMKyle MorenoYCYucheng Chu

Key Points

  • This study aims to evaluate the effectiveness of a multi-agent large language model in automating the coding of pedagogical content knowledge (PCK) in mathematics.
  • Developed a multi-agent LLM, GradeOpt, for coding both content knowledge (CK) and pedagogical content knowledge (PCK).
  • Compared GradeOpt with existing NLP models (RoBERTa, SBERT) and a widely used LLM (GPT-4o).
  • Analyzed responses from a national sample of 268 U.S. middle school mathematics teachers on ratios and proportional reasoning.
  • GradeOpt achieved substantial agreement (κ = 0.68, weighted κ = 0.79) in PCK coding.
  • Outperformed comparison models which had lower agreement (κ ≤ 0.39, weighted κ ≤ 0.55).
  • Demonstrated potential for automated coding to be on par with human coders.

Abstract

Abstract Despite the importance of content-specific knowledge for teaching, the task of measuring such knowledge, particularly pedagogical content knowledge (PCK), remains challenging. Scholars have been more successful in capturing the distinct nature of PCK when it has been assessed through open-ended response items, and such responses have been significantly linked to the quality of mathematics instruction and student learning. However, the use of open-ended items entails substantial resource and time demands, as it requires the training and ongoing calibration of human coders to ensure reliable coding of responses. Although scholars have explored technological approaches to automating coding, traditional automated methods have not achieved sufficient reliability for coding complex constructs, such as PCK. Advances in large language models (LLMs) offer new possibilities; however, it remains unclear whether existing LLMs are adequate or whether a purposefully designed LLM framework is required. In this study, we explored the potential of a multi-agent LLM, GradeOpt, which we developed to code mathematics content knowledge (CK) and PCK reliably, in comparison with commonly used automated coding approaches, including two natural language processing (NLP) models (RoBERTa and SBERT) and a widely used LLM (GPT-4o). Using data collected from a national sample of 268 U. S. middle school mathematics teachers who responded to a set of CK and PCK items on ratios and proportional reasoning, we found that the multi-agent LLM, GradeOpt, achieved substantial agreement overall (=. 68 κ =. 68 ; weighted =. 79 κ =. 79) and substantially outperformed the comparison models (. 39 κ ≤. 39 ; weighted. 55 κ ≤. 55) on PCK items. These results demonstrate that a multi-agent LLM framework has the potential to code open-ended responses at a level comparable to human coders. We discuss the implications for advancing automated assessment in teacher education and mathematics education.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Copur-Gencturk et al. (2026) studied this question.

synapsesocial.com/papers/69fecfafb9154b0b82876a05https://doi.org/10.1007/s11858-026-01796-2
Ask AI
Helpful
Bookmark
Share
View Full Paper