PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 1, 2023141 citations

Analyzing Leakage of Personally Identifiable Information in Language Models

View Full Paper
NLNils LukasASAhmed SalemRSRobert B. Sim

Key Points

Key points are not available for this paper at this time.

Abstract

Language Models (LMs) have been shown to leak information about training data through sentence-level membership inference and reconstruction attacks. Understanding the risk of LMs leaking Personally Identifiable Information (PII) has received less attention, which can be attributed to the false assumption that dataset curation techniques such as scrubbing are sufficient to prevent PII leakage. Scrubbing techniques reduce but do not prevent the risk of PII leakage: in practice scrubbing is imperfect and must balance the trade-off between minimizing disclosure and preserving the utility of the dataset. On the other hand, it is unclear to which extent algorithmic defenses such as differential privacy, designed to guarantee sentence-or user-level privacy, prevent PII disclosure. In this work, we introduce rigorous game-based definitions for three types of PII leakage via black-box extraction, inference, and reconstruction attacks with only API access to an LM. We empirically evaluate the attacks against GPT-2 models fine-tuned with and without defenses in three domains: case law, health care, and e-mails. Our main contributions are (i) novel attacks that can extract up to 10× more PII sequences than existing attacks, (ii) showing that sentence-level differential privacy reduces the risk of PII disclosure but still leaks about 3% of PII sequences, and (iii) a subtle connection between record-level membership inference and PII reconstruction. Code to reproduce all experiments in the paper is available at https: //github. com/microsoft/analysingₚiiₗeakage.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lukas et al. (2023) studied this question.

synapsesocial.com/papers/69d95e8dc7f0c3ae80a3d28chttps://doi.org/10.1109/sp46215.2023.10179300
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 15th International Conference on Learning Representations (ICLR 17)2017 · 3,522 citations
  2. 2Revisiting the uniqueness of simple demographics in the US population2006 · 259 citations
  3. 3A survey of named entity recognition and classification2007 · 2,503 citations