PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 22, 20243 citationsOpen Access

Watermarking Makes Language Models Radioactive

View Full Paper
TSTom SanderPFPierre FernandezADAlain Durmus

Key Points

Key points are not available for this paper at this time.

Abstract

This paper investigates the radioactivity of LLM-generated texts, i.e. whether it is possible to detect that such input was used as training data. Conventional methods like membership inference can carry out this detection with some level of accuracy. We show that watermarked training data leaves traces easier to detect and much more reliable than membership inference. We link the contamination level to the watermark robustness, its proportion in the training set, and the fine-tuning process. We notably demonstrate that training on watermarked synthetic instructions can be detected with high confidence (p-value < 1e-5) even when as little as 5% of training text is watermarked. Thus, LLM watermarking, originally designed for detecting machine-generated text, gives the ability to easily identify if the outputs of a watermarked LLM were used to fine-tune another LLM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sander et al. (2024) studied this question.

synapsesocial.com/papers/68e780d5b6db6435876f406ehttps://doi.org/10.48550/arxiv.2402.14904
Ask AI
Helpful
Bookmark
Share
View Full Paper