PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 21, 20250 citationsOpen Access

SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models

View Full Paper
HCHuijuan CuiWLWei Liu

Key Points

  • This research presents a benchmark aimed at evaluating the resilience of large language models against adversarial prompts.
  • Introduces SecReEvalBench with metrics for assessing resilience to attacks.
  • Evaluates models using a dataset with neutral and malicious prompts.
  • Analyzes five large language models: Llama 3.1, Gemma 2, Mistral v0.3, DeepSeek-R1, and Qwen 3.
  • Provides insights into the effectiveness of large language models against various adversarial strategies.
  • Identifies strengths and weaknesses in model defenses against evolving threats.

Abstract

The increasing deployment of large language models in security-sensitive domains necessitates rigorous evaluation of their resilience against adversarial prompt-based attacks. While previous benchmarks have focused on security evaluations with limited and predefined attack domains, such as cybersecurity attacks, they often lack a comprehensive assessment of intent-driven adversarial prompts and the consideration of real-life scenario-based multi-turn attacks. To address this gap, we present SecReEvalBench, the Security Resilience Evaluation Benchmark, which defines four novel metrics: Prompt Attack Resilience Score, Prompt Attack Refusal Logic Score, Chain-Based Attack Resilience Score and Chain-Based Attack Rejection Time Score. Moreover, SecReEvalBench employs six questioning sequences for model assessment: one-off attack, successive attack, successive reverse attack, alternative attack, sequential ascending attack with escalating threat levels and sequential descending attack with diminishing threat levels. In addition, we introduce a dataset customized for the benchmark, which incorporates both neutral and malicious prompts, categorised across seven security domains and sixteen attack techniques. In applying this benchmark, we systematically evaluate five state-of-the-art open-weighted large language models, Llama 3.1, Gemma 2, Mistral v0.3, DeepSeek-R1 and Qwen 3. Our findings offer critical insights into the strengths and weaknesses of modern large language models in defending against evolving adversarial threats. The SecReEvalBench dataset is publicly available at https://kaggle.com/datasets/5a7ee22cf9dab6c93b55a73f630f6c9b42e936351b0ae98fbae6ddaca7fe248d, which provides a groundwork for advancing research in large language model security.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cui et al. (2025) studied this question.

synapsesocial.com/papers/69473b64db9c958d0dfca951https://doi.org/10.48550/arxiv.2505.07584
Ask AI
Helpful
Bookmark
Share
View Full Paper