PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 20244 citationsOpen Access

Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

View Full Paper
CSChristian SchlarmannNSNaman Deep SinghFCFrancesco Croce

Key Points

Key points are not available for this paper at this time.

Abstract

Multi-modal foundation models like OpenFlamingo, LLaVA, and GPT-4 are increasingly used for various real-world tasks. Prior work has shown that these models are highly vulnerable to adversarial attacks on the vision modality. These attacks can be leveraged to spread fake information or defraud users, and thus pose a significant risk, which makes the robustness of large multi-modal foundation models a pressing problem. The CLIP model, or one of its variants, is used as a frozen vision encoder in many vision-language models (VLMs), e.g. LLaVA and OpenFlamingo. We propose an unsupervised adversarial fine-tuning scheme to obtain a robust CLIP vision encoder, which yields robustness on all vision down-stream tasks (VLMs, zero-shot classification) that rely on CLIP. In particular, we show that stealth-attacks on users of VLMs by a malicious third party providing manipulated images are no longer possible once one replaces the original CLIP model with our robust one. No retraining or fine-tuning of the VLM is required. The code and robust models are available at https://github.com/chs20/RobustVLM

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Schlarmann et al. (2024) studied this question.

synapsesocial.com/papers/68e78950b6db6435876fb7f9https://doi.org/10.48550/arxiv.2402.12336
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective2024
  2. 2Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models2025
  3. 3Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models2025
  4. 4microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification2026
  5. 5Understanding Modality-Specific Vulnerabilities in Vision–Language Models Under Adversarial Attacks2026