Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 15, 2026PLoS ONEOpen Access

A comparative analysis of topic modelling techniques for the thematic analysis of student feedback

View Full Paper
Ask AI
Bookmark
Share

Authors

NKNeha KardamDWDenise Wilson

Discussion

Loading...

Member takes

Overview

Comparative analysis reveals Non-Negative Matrix Factorization outperforms other models in student feedback analysis, indicating the critical role of human domain experts in thematic NLP.

Key Points

  • To identify best practices and evaluate the performance of unsupervised short text topic modeling techniques using natural language processing on semi-structured student feedback data in education research.
  • Analyzed student survey responses gathered between 2016 and 2023 across >40 engineering courses, yielding datasets for faculty support (N=1,667), teaching assistant (TA) support (N=1,592), and peer support (N=1,376).
  • Evaluated five unsupervised models (LDA, LSA, NMF, k-means, and BERTopic) using internal coherence metrics and two ground truth strategies: machine-led manual coding and human-led domain expert independent coding.
  • Non-Negative Matrix Factorization (NMF) achieved the highest overall performance in two datasets, reaching 75.6% accuracy, 75.7% F1-score, and 0.63 interrater reliability for peer support, and 72.6% accuracy, 72.0% F1-score, and 0.57 interrater reliability for TA support.
  • The human-led approach produced higher accuracy and F1-scores for faculty and peer support data, but encountered misalignment with automated topic models in the TA support dataset.

Cite This Study

Kardam et al. (2026) studied this question.

synapsesocial.com/papers/6a8019bb75c2e31742c85d49https://doi.org/10.1371/journal.pone.0328697
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Probabilistic Extension of Precision, Recall, and F1 Score for More Thorough Evaluation of Classification Models2020 · 761 citations
  2. 2Computing Inter-Rater Reliability for Observational Data: An Overview and Tutorial2012 · 3,938 citations
  3. 3Latent semantic analysis2004 · 1,121 citations
  4. 4An Exploratory Analysis of GSDMM and BERTopic on Short Text Topic Modelling2022 · 12 citations