PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 19, 202410 citationsOpen Access

CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

View Full Paper
MBManish BhattSCSahana ChennabasappaYLYue Li

Key Points

Key points are not available for this paper at this time.

Abstract

Large language models (LLMs) introduce new security risks, but there are few comprehensive evaluation suites to measure and reduce these risks. We present BenchmarkName, a novel benchmark to quantify LLM security risks and capabilities. We introduce two new areas for testing: prompt injection and code interpreter abuse. We evaluated multiple state-of-the-art (SOTA) LLMs, including GPT-4, Mistral, Meta Llama 3 70B-Instruct, and Code Llama. Our results show that conditioning away risk of attack remains an unsolved problem; for example, all tested models showed between 26% and 41% successful prompt injection tests. We further introduce the safety-utility tradeoff: conditioning an LLM to reject unsafe prompts can cause the LLM to falsely reject answering benign prompts, which lowers utility. We propose quantifying this tradeoff using False Refusal Rate (FRR). As an illustration, we introduce a novel test set to quantify FRR for cyberattack helpfulness risk. We find many LLMs able to successfully comply with "borderline" benign requests while still rejecting most unsafe requests. Finally, we quantify the utility of LLMs for automating a core cybersecurity task, that of exploiting software vulnerabilities. This is important because the offensive capabilities of LLMs are of intense interest; we quantify this by creating novel test sets for four representative problems. We find that models with coding capabilities perform better than those without, but that further work is needed for LLMs to become proficient at exploit generation. Our code is open source and can be used to evaluate other LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bhatt et al. (2024) studied this question.

synapsesocial.com/papers/68e6e657b6db64358766154ehttps://doi.org/10.48550/arxiv.2404.13161
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1CS-Eval—A Concise Benchmark for Evaluating the Security Risks of Large Language Models2024 · 1 citations
  2. 2SECURE: Benchmarking Generative Large Language Models for Cybersecurity Advisory2024 · 2 citations
  3. 3Evaluating the Cybersecurity Robustness of Commercial LLMs against Adversarial Prompts: A PromptBench Analysis2024 · 3 citations
  4. 4S-Eval: Automatic and Adaptive Test Generation for Benchmarking Safety Evaluation of Large Language Models2024 · 1 citations
  5. 5Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data2025