PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 6, 2026Software Practice and Experience0 citations

Security in the Fine‐Tuning Lifecycle of Large Language Models: Threats, Defenses, Evaluation, and Future Directions

View Full Paper
WLWenjuan LiYLYitao LiuRCRunze Chen

Key Points

  • This paper aims to survey security issues related to fine-tuning large language models and propose a unified framework.
  • Attacks and defenses are categorized into pre-tuning, during-tuning, and post-tuning phases.
  • Evaluation of representative methods using a unified model selection and experimental setup for cross-phase testing.
  • Comparative analysis of attack and defense strategies to understand their relationships and effectiveness.
  • Attack effectiveness varies significantly with model type and size; older weight-editing attacks are less effective on modern models.
  • Cross-lingual backdoor transfer achieves perfect results in larger models but fails on 1B-4B models tested.
  • Single-phase defenses typically do not generalize across different attack phases, indicating a need for robust cross-phase solutions.

Abstract

ABSTRACT Background Fine‐tuning has become a core mechanism for adapting pretrained Large Language Models (LLMs) to downstream tasks. However, the dependence of fine‐tuning on training data, parameter update mechanisms, and reusable components provides entry points for attackers. Related threats have evolved from data poisoning and weight tampering to agent behavior manipulation and interface exploitation, while defensive research has expanded from preimmunization to post‐hoc remediation. However, existing reviews do not yet provide a unified framework that organizes attack and defense methods across the complete fine‐tuning lifecycle. Objective This paper serves as a systematic survey of LLM security in fine‐tuning scenarios and establishes a lifecycle‐based framework for comparing attack and defense methods, complemented by unified empirical evaluation. Methods Fine‐tuning‐related attack and defense mechanisms are divided into three phases according to the timing of intervention: the pre‐tuning, during‐tuning, and post‐tuning phases. Within each phase, attack and defense strategies are reviewed and contrasted to expose their evolutionary relationships and limitations. Representative methods from each phase are then evaluated under a unified model selection, hardware setup, and evaluation protocol, with additional cross‐phase experiments pairing attacks and defenses from different phases. Results Unified evaluation reveals that attack effectiveness is highly model‐dependent and nonmonotonic with scale: weight‐editing attacks that succeed on earlier models lose impact on modern open‐source LLMs; cross‐lingual backdoor transfer, reported as near‐perfect at larger scales, fails entirely on tested 1B‐4B models; and purely benign fine‐tuning samples can compromise safety alignment in instruction‐tuned models. Cross‐phase experiments further show that single‐phase defenses rarely generalize to attacks from other phases, and that defense effectiveness depends jointly on model architecture and alignment state. Conclusion Based on the survey and experimental findings, this paper identifies key open problems, including configuration‐robust defense, cross‐phase defense composition, and embedding‐space attacks beyond behavioral assumptions‐and proposes concrete directions for future research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2026) studied this question.

synapsesocial.com/papers/6a743783764cddc9499d4e1fhttps://doi.org/10.1002/spe.70098
Ask AI
Helpful
Bookmark
Share
View Full Paper