PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 19, 20240 citationsOpen Access

Transferring Backdoors between Large Language Models by Knowledge Distillation

View Full Paper
PCPengzhou ChengZWZongru WuTJTianjie Ju

Key Points

Key points are not available for this paper at this time.

Abstract

Backdoor Attacks have been a serious vulnerability against Large Language Models (LLMs). However, previous methods only reveal such risk in specific models, or present tasks transferability after attacking the pre-trained phase. So, how risky is the model transferability of a backdoor attack? In this paper, we focus on whether existing mini-LLMs may be unconsciously instructed in backdoor knowledge by poisoned teacher LLMs through knowledge distillation (KD). Specifically, we propose ATBA, an adaptive transferable backdoor attack, which can effectively distill the backdoor of teacher LLMs into small models when only executing clean-tuning. We first propose the Target Trigger Generation (TTG) module that filters out a set of indicative trigger candidates from the token list based on cosine similarity distribution. Then, we exploit a shadow model to imitate the distilling process and introduce an Adaptive Trigger Optimization (ATO) module to realize a gradient-based greedy feedback to search optimal triggers. Extensive experiments show that ATBA generates not only positive guidance for student models but also implicitly transfers backdoor knowledge. Our attack is robust and stealthy, with over 80% backdoor transferability, and hopes the attention of security.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cheng et al. (2024) studied this question.

synapsesocial.com/papers/68e5bd3ab6db643587554ecahttps://doi.org/10.48550/arxiv.2408.09878
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning2024 · 2 citations
  2. 2Data Stealing Attacks against Large Language Models via Backdooring2024 · 11 citations
  3. 3Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution2025
  4. 4Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack2025
  5. 5BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models2024 · 2 citations