Synthetic data plays a crucial role in augmenting datasets for machine learning, particularly when real data is limited or sensitive. However, not all synthetic samples contribute positively to model performance; low-quality instances can degrade both predictive accuracy and data fidelity. To address this challenge, we propose the Quality Data Extractor (QDE), a novel framework for filtering synthetic data before augmentation. QDE offers two complementary techniques. The Comprehensive Extraction Strategy (CES) retains synthetic samples that improve classification performance. At the same time, the Optimal Extraction Strategy (OES) combines classifier accuracy with feature-space distance to identify and remove redundant or harmful instances. Together, these strategies enable flexible and robust data filtration across different synthetic data generators and learning tasks. We evaluate QDE on three benchmark datasets—Epileptic Seizure Recognition (ESR), WUSTL Human Activity, and German Credit Card (GCC) dataset using synthetic data generated by CTGAN, Copula GAN, TVAE, and Gaussian Copula. The performance of filtered datasets is evaluated across various statistical and predictive metrics, including the KL divergence, Anderson-Darling statistic, mean difference, propensity score range, and standard classification metrics. Experiments are conducted using Gaussian Naive Bayes (GaussianNB) and XGBoost (XGB) classifiers. Results show that both CES and OES consistently outperform naive augmentation. OES is most effective in low-quality synthetic scenarios, improving classification accuracy by up to 12.4% and reducing KL divergence by over 95% in low-quality synthetic data scenarios. CES excels when synthetic data already exhibits strong predictive performance, offering reliable gains with lower computational overhead. These findings underscore the need for post-generation filtering and position QDE as a practical and extensible solution for integrating high-quality synthetic data. QDE achieves this without requiring changes to generation models, making it a drop-in filtration step in existing pipelines.
No takes yet. Share an insight, caveat, or question.
Sachdeva et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: