PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 3, 20250 citationsOpen Access

Transferable Adversarial Attacks on Black-Box Vision-Language Models

View Full Paper
KHKai HuState Key Laboratory of Vehicle NVH and Safety TechnologyWYWeichen YuCommercial Aircraft Corporation of China (China)LZLi ZhangNanjing University of Chinese Medicine

Key Points

  • Targeted adversarial examples transfer effectively to proprietary vision-language models, including GPT-4o.
  • Universal perturbations can misinterpret visual information across multiple proprietary VLLMs.
  • Significant vulnerabilities in object recognition and visual question answering highlight the risks of current models.
  • There is an urgent requirement for strong mitigation measures to protect users from these attacks.

Abstract

Vision Large Language Models (VLLMs) are increasingly deployed to offer advanced capabilities on inputs comprising both text and images. While prior research has shown that adversarial attacks can transfer from open-source to proprietary black-box models in text-only and vision-only contexts, the extent and effectiveness of such vulnerabilities remain underexplored for VLLMs. We present a comprehensive analysis demonstrating that targeted adversarial examples are highly transferable to widely-used proprietary VLLMs such as GPT-4o, Claude, and Gemini. We show that attackers can craft perturbations to induce specific attacker-chosen interpretations of visual information, such as misinterpreting hazardous content as safe, overlooking sensitive or restricted material, or generating detailed incorrect responses aligned with the attacker's intent. Furthermore, we discover that universal perturbations -- modifications applicable to a wide set of images -- can consistently induce these misinterpretations across multiple proprietary VLLMs. Our experimental results on object recognition, visual question answering, and image captioning show that this vulnerability is common across current state-of-the-art models, and underscore an urgent need for robust mitigations to ensure the safe and secure deployment of VLLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hu et al. (2025) studied this question.

synapsesocial.com/papers/68e03501f0e39f13e7fa39e3https://doi.org/10.48550/arxiv.2505.01050
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Attention! You Vision Language Model Could Be Maliciously Manipulated2025
  2. 2Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models2025 · 2 citations
  3. 3White-box Multimodal Jailbreaks Against Large Vision-Language Models2024
  4. 4Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack2026
  5. 5VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models2025