PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs

View Full Paper
FRFatmazohra RezkellahRDRamzi Dakhmouche

Key Points

  • The proposed method improves adversarial robustness while effectively unlearning sensitive information from LLMs.
  • Using point-wise constraint-based interventions yields superior performance compared to existing methods with lower computational cost.
  • The research highlights the effectiveness of constrained optimization in handling both unlearning and robustness without needing an oracle classifier.
  • Findings indicate that simpler optimization can outperform more complex max-min interventions in maintaining model integrity.

Abstract

With the increasing adoption of Large Language Models (LLMs), more customization is needed to ensure privacy-preserving and safe generation. We address this objective from two critical aspects: unlearning of sensitive information and robustness to jail-breaking attacks. We investigate various constrained optimization formulations that address both aspects in a unified manner, by finding the smallest possible interventions on LLM weights that either make a given vocabulary set unreachable or embed the LLM with robustness to tailored attacks by shifting part of the weights to a safer region. Beyond unifying two key properties, this approach contrasts with previous work in that it doesn't require an oracle classifier that is typically not available or represents a computational overhead. Surprisingly, we find that the simplest point-wise constraint-based intervention we propose leads to better performance than max-min interventions, while having a lower computational cost. Comparison against state-of-the-art defense methods demonstrates superior performance of the proposed approach.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rezkellah et al. (2025) studied this question.

synapsesocial.com/papers/68e865117ef2f04ca37e4cechttps://doi.org/10.48550/arxiv.2510.03567
Ask AI
Helpful
Bookmark
Share
View Full Paper