PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 3, 20250 citationsOpen Access

SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection

View Full Paper
MJMaithili JoshiPNPalash NandiTCTanmoy Chakraborty

Key Points

  • The SABER method improves safety alignment mechanisms in large language models by utilizing a unique residual connection.
  • Using this approach, we achieved a 51% performance improvement on the HarmBench test set compared to previous methods.
  • Our evaluation shows that SABER has only a slight impact on perplexity when tested on the HarmBench validation set.
  • These findings highlight the importance of addressing vulnerabilities in safety alignment to mitigate jailbreak attacks effectively.

Abstract

Large Language Models (LLMs) with safe-alignment training are powerful instruments with robust language comprehension capabilities. These models typically undergo meticulous alignment procedures involving human feedback to ensure the acceptance of safe inputs while rejecting harmful or unsafe ones. However, despite their massive scale and alignment efforts, LLMs remain vulnerable to jailbreak attacks, where malicious users manipulate the model to produce harmful outputs that it was explicitly trained to avoid. In this study, we find that the safety mechanisms in LLMs are predominantly embedded in the middle-to-late layers. Building on this insight, we introduce a novel white-box jailbreak method, SABER (Safety Alignment Bypass via Extra Residuals), which connects two intermediate layers s and e such that s < e, through a residual connection. Our approach achieves a 51% improvement over the best-performing baseline on the HarmBench test set. Furthermore, SABER induces only a marginal shift in perplexity when evaluated on the HarmBench validation set. The source code is publicly available at https: //github. com/PalGitts/SABER.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Joshi et al. (2025) studied this question.

synapsesocial.com/papers/68e040eda99c246f578b33e4https://doi.org/10.48550/arxiv.2509.16060
Ask AI
Helpful
Bookmark
Share
View Full Paper