PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 25, 2026Journal of King Abdulaziz University-Engineering Sciences0 citationsOpen Access

Assessing the Use of Synthetic Tabular Data: An Analysis of the Adult Dataset

View Full Paper
ABAbdullah Balamash

Key Points

  • Evaluate the effectiveness of synthetic data generated by different models in comparison to real data.
  • Evaluate real data as a baseline.
  • Assess synthetic data's predictive features against real data in various training scenarios.
  • Implement minority class oversampling to analyze performance in class-imbalanced situations.
  • Synthetic data retains predictive features of baseline data, indicating its potential use.
  • Augmentation of real data does not enhance performance with sufficient existing data.
  • Minority class oversampling using CTGAN improves recall from 0.60 to 0.80 but lowers precision, underperforming compared to traditional methods.

Abstract

In this study, we evaluate the effectiveness of synthetic data generated using the Gaussian Copula (GC) and Conditional Tabular Generative Adversarial Network (CTGAN) models within the Synthetic Data Vault (SDV) framework by applying it to the Adult data set 1. This is done in three steps:(1) Training and Testing on Real Data (Baseline), (2) Training on Synthetic Data and Testing on Real Data (TSTR), (3) Training on (Synthetic + Real Data) and Testing on Real Data (Augmentation with Real Data), (4) Minority Class Oversampling. The TSTR results show that synthetic data preserves the predictive features of baseline data. Augmentation with real data does not improve the performance when there is enough real data. When we have a class-imbalance scenario, synthetic minority oversampling improves the recall for the minority class (e.g., from 0.60 to 0.80 with CTGAN) at the expense of the precision in a way that underperforms traditional techniques such as random oversampling and class weighting. Overall, our findings suggest that synthetic data can be used when we do not have enough data, but it is not good enough to address class imbalance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abdullah Balamash (2026) studied this question.

synapsesocial.com/papers/6a13e7cf0e02ee3982d32784https://doi.org/10.64064/1658-4260.1020
Ask AI
Helpful
Bookmark
Share
View Full Paper