PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 6, 2026Electronics0 citationsOpen Access

Jailbreaking MLLMs via Attention Redirection and Entropy Regularization

View Full Paper
JDJiayu DuFDFangxu DongFZFan Zhang

Key Points

  • This research aims to enhance the efficacy of jailbreak attacks on multimodal large language models by addressing safety vulnerabilities.
  • Introduces Attention-Enhancement and Targeted Entropy Regularization for Adversarial Optimization (AERO).
  • Develops an attention enhancement loss to redirect attention towards visual tokens.
  • Implements a targeted entropy regularization scheme to encourage output diversity.
  • Achieves Attack Success Rates (ASRs) of 65.8–70.7% on MM-SafetyBench and 71.0–84.5% on HarmBench.
  • Surpasses strongest baselines by up to 16.2% in success rate.
  • Generates higher-quality harmful content consistently.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across vision–language tasks, yet their safety alignment remains vulnerable to adversarial manipulation. Existing jailbreak attacks typically optimize adversarial perturbations using negative log-likelihood loss alone, which often leads to overfitting on target affirmative tokens and fails to elicit substantive harmful content. We propose Attention-Enhancement and Targeted Entropy Regularization for Adversarial Optimization (AERO), a novel jailbreak framework addressing these limitations through two complementary mechanisms. First, an attention enhancement loss strategically redirects cross-modal attention toward perturbed visual tokens, distracting safety-aligned features from scrutinizing malicious queries. Second, a targeted entropy regularization scheme maximizes output diversity over non-refusal tokens during initial generation, creating a permissive context that improves cross-query generalization and enables responses that genuinely address malicious requests. Extensive experiments on multiple state-of-the-art MLLMs demonstrate that AERO significantly outperforms existing methods, achieving Attack Success Rates (ASRs) of 65.8–70.7% on MM-SafetyBench and 71.0–84.5% on HarmBench. Our approach surpasses the strongest baselines by margins of up to 16.2% in success rate while consistently generating higher-quality harmful content.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Du et al. (2026) studied this question.

synapsesocial.com/papers/695d85653483e917927a4f3ehttps://doi.org/10.3390/electronics15010237
Ask AI
Helpful
Bookmark
Share
View Full Paper