PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 19, 2025IEEE Transactions on Image Processing0 citations

The Safety Illusion? Testing the Boundaries of Concept Removal in Diffusion Models

View Full Paper
YPYutian PanTLTing LuoYLY. Li

Key Points

  • Concept erasure methods are ineffective at fully eliminating sensitive concepts, especially NSFW content.
  • A novel attack framework, CEA, significantly enhances the adversarial success rate in regenerating erased concepts.
  • Existing techniques struggle against adversarial prompts targeting erased concepts, revealing security flaws.
  • Extensive experiments demonstrate that CEA exploits vulnerabilities in generative latent space for effective concept regeneration.

Abstract

Text-to-image diffusion models are capable of producing high-quality images from textual descriptions; however, they present notable security concerns. These include the potential for generating Not-Safe-For-Work (NSFW) content, replicating artists' styles without authorization, or creating deepfakes. Recent advancements have proposed concept erasure techniques to eliminate sensitive concepts from these models, aiming to mitigate the generation of undesirable content. Nevertheless, the robustness of these techniques against a wide range of adversarial inputs has not been comprehensively investigated. To address this challenge, a novel two-stage optimization attack framework based on adversarial perturbations, referred to as Concept Embedding Adversary (CEA), was proposed in the present study. By leveraging the cross-modal alignment priors of the CLIP model, CEA iteratively adjusts adversarial embedding vectors to approximate the semantic expression of specific target concepts. This process enables the construction of deceptive adversarial prompts that exploit diffusion models, compelling them to regenerate previously erased concepts. The performance of concept erasure methods was evaluated, specifically when dealing with diversified adversarial prompts targeting erased concepts, such as NSFW content, artistic styles, and objects. Extensive experimental results demonstrate that existing concept erasure methods are unable to completely eliminate target concepts. In contrast, the proposed CEA framework exploits residual vulnerabilities within the generative latent space through a two-stage optimization process. By achieving precise cross-modal alignment, CEA attains significantly higher ASR in regenerating erased concepts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pan et al. (2025) studied this question.

synapsesocial.com/papers/68f500b442a2eee15b0a0dfahttps://doi.org/10.1109/tip.2025.3620665
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ImageNet: A large-scale hierarchical image database2009 · 63,148 citations
  2. 2MMA-Diffusion: MultiModal Attack on Diffusion Models2024 · 48 citations
  3. 3Perception-Guided Jailbreak Against Text-to-Image Models2025 · 11 citations
  4. 4Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient2025 · 12 citations
  5. 5Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models2023 · 186 citations