PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 28, 202317 citationsOpen Access

Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4

KCKent K. ChangMCMackenzie Hạnh CramerSSSandeep Soni

Key Points

  • This work investigates the extent to which ChatGPT and GPT-4 have memorized copyrighted books based on online frequency.
  • Conducted a name cloze membership inference query to identify memorized books.
  • Analyzed memorization frequency in relation to web availability of text passages from books.
  • Compared model performance on memorized versus non-memorized texts in downstream tasks.
  • ChatGPT and GPT-4 memorized a substantial range of copyrighted materials.
  • Models showed significantly better performance on memorized books compared to non-memorized ones.
  • Findings suggest implications for the validity of cultural analytics tests due to data contamination.

Abstract

In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query. We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with which passages of those books appear on the web. The ability of these models to memorize an unknown set of books complicates assessments of measurement validity for cultural analytics by contaminating test data; we show that models perform much better on memorized books than on non-memorized books for downstream tasks. We argue that this supports a case for open models whose training data is known.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chang et al. (2023) studied this question.

synapsesocial.com/papers/6a0e9a73f59e0974004c461bhttps://doi.org/10.48550/arxiv.2305.00118
Ask AI
Helpful
Bookmark
Share
View Full Paper