PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 13, 20250 citationsOpen Access

AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses

View Full Paper
NCNicholas CarliniJRJavier RandoEDEdoardo Debenedetti

Key Points

  • AutoAdvExBench evaluates LLMs in adversarial example defense tasks, with 75% success on CTF-like defenses.
  • The designed agent attacks 13% of real-world defenses, highlighting a significant gap between CTF-like defenses and real ones.
  • Stronger LLMs succeed on 21% of real defenses but show only 54% success on CTF-like defenses.
  • The benchmark aims to enhance practical applications for adversarial machine learning research.

Abstract

We introduce AutoAdvExBench, a benchmark to evaluate if large language models (LLMs) can autonomously exploit defenses to adversarial examples. Unlike existing security benchmarks that often serve as proxies for real-world tasks, bench directly measures LLMs' success on tasks regularly performed by machine learning security experts. This approach offers a significant advantage: if a LLM could solve the challenges presented in bench, it would immediately present practical utility for adversarial machine learning researchers. We then design a strong agent that is capable of breaking 75% of CTF-like ("homework exercise") adversarial example defenses. However, we show that this agent is only able to succeed on 13% of the real-world defenses in our benchmark, indicating the large gap between difficulty in attacking "real" code, and CTF-like code. In contrast, a stronger LLM that can attack 21% of real defenses only succeeds on 54% of CTF-like defenses. We make this benchmark available at https://github.com/ethz-spylab/AutoAdvExBench.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Carlini et al. (2025) studied this question.

synapsesocial.com/papers/68ece2abd1bb2827d1297460https://doi.org/10.48550/arxiv.2503.01811
Ask AI
Helpful
Bookmark
Share
View Full Paper