PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 6, 2026Algorithms2 citationsOpen Access

Benchmarking Statistical and Deep Generative Models for Privacy-Preserving Synthetic Student Data in Educational Data Mining

View Full Paper
GKGeorgios KostopoulosMTMaria TsiakmakiSKSotiris Kotsiantis

Key Points

  • This research aims to benchmark generative models for creating synthetic student data while ensuring privacy.
  • Benchmark four generative approaches: Gaussian Copula, CopulaGAN, CTGAN, and TVAE.
  • Use the Synthetic Data Vault framework for synthetic data generation.
  • Evaluate fidelity and machine learning utility using Random Forest classifiers.
  • Synthetic data achieved 96-98% predictive performance compared to real data.
  • TVAE showed the highest multivariate fidelity among the generative models.

Abstract

Educational Data Mining (EDM) increasingly depends on large, high-quality datasets to drive predictive and adaptive learning systems. However, data scarcity, privacy restrictions, and limited accessibility severely hinder research reproducibility and cross-institutional collaboration. Synthetic data generation provides an emerging solution, enabling the creation of artificial yet statistically realistic datasets that preserve analytical utility while preserving student privacy. This study benchmarks four generative approaches, namely Gaussian Copula, CopulaGAN, Conditional Tabular Generative Adversarial Networks (CTGAN), and Tabular Variational Auto Encoders (TVAE), on student data from six undergraduate courses at a European university. Using the open-source Synthetic Data Vault (SDV) framework, we evaluate the fidelity and Machine Learning utility of synthetic student records through Random Forest classifiers across five metrics, namely accuracy, F1-score, precision, recall, and Area Under Curve (AUC). The results show that synthetic data can achieve 96–98% of the predictive performance obtained when training on real data, with TVAE consistently demonstrating the highest multivariate fidelity. Our contributions are threefold: (i) we introduce a reproducible benchmarking pipeline for synthetic data evaluation in educational settings; (ii) we empirically compare statistical and deep generative synthesizers on real-world tabular student data; and (iii) we identify critical research directions related to privacy and reproducibility. The findings position synthetic data generation as a foundational technology for ethical and privacy-preserving EDM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kostopoulos et al. (2026) studied this question.

synapsesocial.com/papers/695d856e3483e917927a52d2https://doi.org/10.3390/a19010039
Ask AI
Helpful
Bookmark
Share
View Full Paper