PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 24, 202130 citationsOpen Access

Counterfactual Memorization in Neural Language Models

CZChiyuan ZhangDIDaphne IppolitoKLKatherine Lee

Key Points

Key points are not available for this paper at this time.

Abstract

Modern neural language models that are widely used in various NLP tasks risk memorizing sensitive information from their training data. Understanding this memorization is important in real world applications and also from a learning-theoretical perspective. An open question in previous studies of language model memorization is how to filter out "common" memorization. In fact, most memorization criteria strongly correlate with the number of occurrences in the training set, capturing memorized familiar phrases, public knowledge, templated texts, or other repeated data. We formulate a notion of counterfactual memorization which characterizes how a model's predictions change if a particular document is omitted during training. We identify and study counterfactually-memorized training examples in standard text datasets. We estimate the influence of each memorized training example on the validation set and on generated texts, showing how this can provide direct evidence of the source of memorization at test time.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2021) studied this question.

synapsesocial.com/papers/6a0eab4ffca5c6c9f447a035https://doi.org/10.48550/arxiv.2112.12938
Ask AI
Helpful
Bookmark
Share
View Full Paper