PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 2026Scientific Reports0 citationsOpen Access

Exploring teacher-student interaction through multimodal large language models: an empirical investigation

View Full Paper
GCGuanyu ChenGHGuangxin HanJNJuan Niu

Key Points

  • This research aims to explore the effectiveness of multimodal large language models in analyzing teacher-student interactions.
  • Fine-tuned VisualGLM-6B on 2,380 annotated images from 30 classroom videos.
  • Analyzed five interaction types: guided, collaborative, questioning, independent, and exhibitive.
  • Employed LoRA-based fine-tuning and prompt engineering to improve model accuracy.
  • Assessment included confusion matrices and comparisons with expert annotations.
  • Achieved 82% overall accuracy, especially high on guided, independent, and exhibitive interactions.
  • Collaborative and questioning interactions remained challenging.
  • Provided more structured and interpretable outputs compared to expert annotations, despite some misclassifications and hallucinations.

Abstract

Teacher–student interaction is central to classroom learning, yet traditional observation and machine-learning approaches often remain inefficient and subjective. This study explores the use of multimodal large language models (MLLMs) for systematic analysis of classroom dynamics. We fine-tuned VisualGLM-6B on 2,380 annotated images from 30 classroom videos, covering five interaction types: guided, collaborative, questioning, independent, and exhibitive. LoRA-based fine-tuning combined with prompt engineering was employed to enhance interpretability and domain-specific accuracy. Model performance was assessed through confusion matrices, BERTScore, and expert comparisons. The fine-tuned model achieved 82% overall accuracy, performing best on guided, independent, and exhibitive interactions, while collaborative and questioning types remained challenging. Compared with expert annotation, the model provided more structured and interpretable outputs, though occasional misclassifications and hallucinations persisted. These findings demonstrate the feasibility of applying MLLMs for efficient, objective analysis of teacher–student interactions and highlight future directions such as incorporating audio inputs and larger datasets to further advance educational research methodologies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2026) studied this question.

synapsesocial.com/papers/6987eb5df6bacdd2fe8fc828https://doi.org/10.1038/s41598-026-38626-0
Ask AI
Helpful
Bookmark
Share
View Full Paper