PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 6, 20260 citationsOpen Access

Jailbreaking LLMs: A Survey of Attacks, Defenses and Evaluation

View Full Paper
SHSafayat Bin HakimKGKanchon GharamiNGNahid Farhady Ghalaty

Key Points

  • To systematically analyze security threats posed by jailbreak attacks on large language models and evaluate defensive strategies.
  • Synthesize research from major security and AI venues from 2022 to 2025.
  • Introduce a comprehensive taxonomy of jailbreak vectors.
  • Examine attack success rates across various model types and defense mechanisms.
  • Evaluate the reliability of automated judge agreement for security assessments.
  • Automated attacks achieve 90-99% success on open-weight models.
  • Black-box attacks reach 80-94% effectiveness on proprietary models.
  • Agent-driven attacks achieve 95% success by manipulating conversations across multiple turns.
  • Defensive strategies often fail, with residual success rates above 15% even against protections.

Abstract

Large Language Models (LLMs) excel at natural language understanding and generation, but deployment introduces critical security risks through jailbreak attacks that circumvent safety alignment. This survey provides the first unified rigorous systematization of the LLM security threat landscape (2022-2025), synthesizingresearch across premier security and AI venues. We introduce a comprehensive taxonomy of jailbreak vectors from elementary prompt manipulation to sophisticated multimodal exploits. We find a persistent asymmetry between attack sophistication and defensive capability: advanced automated attacks routinely achieve 90-99% success on open-weight models, while black-box attacks reach 80-94% effectiveness on proprietary models. Agent-driven multi-turn attacks demonstrate 95% success by decomposing harmful queries across conversation turns. Embodied AI vulnerabilities enable jailbreaks to trigger harmful physical actions in robotic platforms, expanding threats beyond digital domains. Defenses show fundamental limits: feedback-based attacks often retain residual success above 15% even against layered protections. Evaluation remains unreliable, with automated judge agreement varying 70-93% depending on implementation, weakening confidence in security assessments. By examining attack transferability, computational costs, and defense overhead, we identify gaps in multimodal safety protocols, agent-based security frameworks, and mechanistic interpretability. Robust LLM security requires shifting from reactive mitigation to proactive security-by-design architectures integrating constitutional AI principles with formal verification.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hakim et al. (2026) studied this question.

synapsesocial.com/papers/698585cb8f7c464f230097dchttps://doi.org/10.13016/m2ixjy-kne6
Ask AI
Helpful
Bookmark
Share
View Full Paper