PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026IEEE Transactions on Pattern Analysis and Machine Intelligence13 citations

Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

View Full Paper
ZWZeming WeiYWYue WangALAng Li

Key Points

  • The research aims to investigate the effectiveness of In-Context Learning in enhancing the safety alignment of language models against malicious exploitation.
  • Proposed In-Context Attack (ICA) and In-Context Defense (ICD) strategies.
  • Utilized minimal in-context demonstrations to manipulate safety responses.
  • Conducted empirical validations across various models, datasets, and attack scenarios.
  • Demonstrated effectiveness of ICA and ICD in altering safety alignment of LLM outputs.
  • Showed scalability of these methods for real-world deployment and red-teaming evaluations.
  • Provided theoretical insights supporting the manipulation of safety through in-context demonstrations.

Abstract

Large Language Models (LLMs) have demonstrated remarkable success across diverse applications, yet their susceptibility to malicious exploitation remains a critical challenge. Notably, LLMs are known to be vulnerable to jailbreaking attacks, where adversaries craft malicious inputs to induce harmful or unethical outputs. In this paper, motivated by the unique effectiveness and scalability of In-Context Learning (ICL) in LLMs, we explore its potential to modulate the safety alignment of LLMs. Specifically, we propose the In-Context Attack (ICA), which employs harmful demonstrations to subvert LLMs' safety, and the In-Context Defense (ICD), which bolsters their resilience through examples that demonstrate refusal to produce harmful responses. By adjusting the distribution of safety in LLM outputs through adversarial demonstrations, our proposed in-context attack and defense facilitate effective manipulation of their alignment. We first provide theoretical insights to illustrate how minimal in-context demonstrations can efficiently alter safety alignment. Empirically, we validate ICA and ICD across multiple models, datasets, and attack baselines, showing their efficacy and scalability for red-teaming evaluations and robust safeguards for real-world deployment. Overall, our work unveils the pivotal yet understudied role of ICL in LLM safety, opening new avenues for understanding and improving them.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wei et al. (2026) studied this question.

synapsesocial.com/papers/698434cff1d9ada3c1fb3605https://doi.org/10.1109/tpami.2026.3660147
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Addressing Prompt Injection in Large Language Models via In-Context Learning2026
  2. 2Safe and Efficient In-Context Learning via Risk Control2025
  3. 3Defending Jailbreak Prompts via In-Context Adversarial Game2024 · 2 citations
  4. 4Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing2024
  5. 5How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States2024