PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 30, 20243 citationsOpen Access

Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks

View Full Paper
XCXiong ChenXQXiangyu QiPCPin‐Yu Chen

Key Points

Key points are not available for this paper at this time.

Abstract

Safety, security, and compliance are essential requirements when aligning large language models (LLMs). However, many seemingly aligned LLMs are soon shown to be susceptible to jailbreak attacks. These attacks aim to circumvent the models' safety guardrails and security mechanisms by introducing jailbreak prompts into malicious queries. In response to these challenges, this paper introduces Defensive Prompt Patch (DPP), a novel prompt-based defense mechanism specifically designed to protect LLMs against such sophisticated jailbreak strategies. Unlike previous approaches, which have often compromised the utility of the model for the sake of safety, DPP is designed to achieve a minimal Attack Success Rate (ASR) while preserving the high utility of LLMs. Our method uses strategically designed interpretable suffix prompts that effectively thwart a wide range of standard and adaptive jailbreak techniques. Empirical results conducted on LLAMA-2-7B-Chat and Mistral-7B-Instruct-v0.2 models demonstrate the robustness and adaptability of DPP, showing significant reductions in ASR with negligible impact on utility. Our approach not only outperforms existing defense strategies in balancing safety and functionality, but also provides a scalable and interpretable solution applicable to various LLM platforms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2024) studied this question.

synapsesocial.com/papers/68e67a9ab6db643587604836https://doi.org/10.48550/arxiv.2405.20099
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing2024
  2. 2LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper2024 · 1 citations
  3. 3Studious Bob Fight Back Against Jailbreaking via Prompt Adversarial Tuning2024
  4. 4Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models2025
  5. 5Jailbreak Attacks and Defenses Against Large Language Models: A Survey2024 · 7 citations