PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 22, 20250 citations

Data-level sampling for dealing with imbalanced datasets: better protection against membership inference attacks

View Full Paper
KSKarla F. C. da SilvaABAntonio de Abreu Batista-JúniorJMJesús Pascual Mena‐Chalco

Key Points

  • Models trained using cost-sensitive learning show increased vulnerability to membership inference attacks.
  • Experiments conducted on the UCI Adult dataset and APS dataset confirm the findings regarding privacy risks.
  • Algorithm-level adjustments in imbalanced datasets reveal more sensitive information during inference.
  • The study highlights the need for better protective strategies against privacy risks in machine learning.

Abstract

The use of machine learning models trained on imbalanced datasets with sensitive information has raised privacy concerns. One significant threat is the Membership Inference Attack (MIA), which tries to figure out if a particular data point was included in the training set. This paper investigates whether algorithm-level cost-sensitive learning poses a greater privacy leakage risk than data-level sampling. We conducted experiments using two datasets: the UCI Adult dataset, which focuses on predicting income, and the APS dataset, which focuses on predicting scientific productivity. Our results indicate that models trained with cost-sensitive learning are more vulnerable to MIAs. This supports the hypothesis that correcting for imbalances at the algorithm level can reveal more private information.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Silva et al. (2025) studied this question.

synapsesocial.com/papers/68f8a381c0c01e5ef8abddbahttps://doi.org/10.5753/sbbd.2025.247478
Ask AI
Helpful
Bookmark
Share
View Full Paper