Computational study demonstrates improved predictive performance under sparse learner interactions, highlighting the value of self-supervised representations in education.
This study proposes a Transformer-based self-supervised learning (SSL) framework for data-efficient educational intelligence under sparse and heterogeneous learner interaction conditions. The framework integrates masked educational interaction reconstruction, contrastive representation learning, temporal interaction-gap modeling, and Transformer-based contextual encoding. Experiments on EdNet used a learner-disjoint training/validation/test protocol to prevent cross-learner information leakage. In the full-label setting, the proposed SSL Transformer achieved an Accuracy of 0.5440, a Macro F1-score of 0.4772, and an AUC of 0.5225. Relative to the supervised Transformer baseline, these results represent improvements of approximately 3.5%, 4.1%, and 6.1%, respectively, although the proposed model did not maximize raw Accuracy across all baselines. With only 10% of the labeled training data, AUC remained 0.5203 versus 0.5225 under full supervision, while Accuracy and Macro F1-score were 0.4064 and 0.3978. Under the most restrictive cold-start condition of 1–3 observed interactions, the model achieved an Accuracy of 0.5063, a Macro F1-score of 0.4585, and an AUC of 0.5066. These findings support the value of self-supervised contextual representation learning under limited-label and sparse-history educational conditions.
No takes yet. Share an insight, caveat, or question.
Chen et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: