PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 20250 citationsOpen Access

Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position

View Full Paper
ZXZhong XieXSX. Y. SongJLJun Luo

Key Points

  • Middle tokens are more critical for safety, enhancing overall model performance and security.
  • MOSA, a new safety alignment method, boosts resilience against attacks by focusing on token generation.
  • Assessment compared the effectiveness of MOSA against eight attack methods on varied benchmarks.
  • Identifying asymmetries between defenders and attackers informs stronger safety mechanisms for language models.

Abstract

Diffusion Large Language Models (dLLMs) have recently emerged as a competitive non-autoregressive paradigm due to their unique training and inference approach. However, there is currently a lack of safety study on this novel architecture. In this paper, we present the first analysis of dLLMs' safety performance and propose a novel safety alignment method tailored to their unique generation characteristics. Specifically, we identify a critical asymmetry between the defender and attacker in terms of security. For the defender, we reveal that the middle tokens of the response, rather than the initial ones, are more critical to the overall safety of dLLM outputs; this seems to suggest that aligning middle tokens can be more beneficial to the defender. The attacker, on the contrary, may have limited power to manipulate middle tokens, as we find dLLMs have a strong tendency towards a sequential generation order in practice, forcing the attack to meet this distribution and diverting it from influencing the critical middle tokens. Building on this asymmetry, we introduce Middle-tOken Safety Alignment (MOSA), a novel method that directly aligns the model's middle generation with safe refusals exploiting reinforcement learning. We implement MOSA and compare its security performance against eight attack methods on two benchmarks. We also test the utility of MOSA-aligned dLLM on coding, math, and general reasoning. The results strongly prove the superiority of MOSA.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xie et al. (2025) studied this question.

synapsesocial.com/papers/68d913a34ddcf71ba560b8d6https://doi.org/10.48550/arxiv.2508.12398
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models2025
  2. 2Safety Alignment Should Be Made More Than Just a Few Tokens Deep2024 · 4 citations
  3. 3Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs2024
  4. 4ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning2025
  5. 5Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!2024