PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 30, 20240 citationsOpen Access

Privacy Backdoors: Stealing Data with Corrupted Pretrained Models

View Full Paper
SFShanglun FengFTFlorian Tramèr

Key Points

Key points are not available for this paper at this time.

Abstract

Practitioners commonly download pretrained machine learning models from open repositories and finetune them to fit specific applications. We show that this practice introduces a new risk of privacy backdoors. By tampering with a pretrained model's weights, an attacker can fully compromise the privacy of the finetuning data. We show how to build privacy backdoors for a variety of models, including transformers, which enable an attacker to reconstruct individual finetuning samples, with a guaranteed success! We further show that backdoored models allow for tight privacy attacks on models trained with differential privacy (DP). The common optimistic practice of training DP models with loose privacy guarantees is thus insecure if the model is not trusted. Overall, our work highlights a crucial and overlooked supply chain attack on machine learning privacy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Feng et al. (2024) studied this question.

synapsesocial.com/papers/68e71abfb6db6435876948d9https://doi.org/10.48550/arxiv.2404.00473
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models2024 · 1 citations
  2. 2PreCurious: How Innocent Pre-Trained Language Models Turn into Privacy Traps2024
  3. 3Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage2024 · 1 citations
  4. 4TMI! Finetuned Models Leak Private Information from their Pretraining Data2024 · 10 citations
  5. 5Data Stealing Attacks against Large Language Models via Backdooring2024 · 11 citations