PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 20242 citationsOpen Access

Uncovering Safety Risks in Open-source LLMs through Concept Activation Vector

View Full Paper
ZXZhihao XuRHRuixuan HuangXWXiting Wang

Key Points

Key points are not available for this paper at this time.

Abstract

Current open-source large language models (LLMs) are often undergone careful safety alignment before public release. Some attack methods have also been proposed that help check for safety vulnerabilities in LLMs to ensure alignment robustness. However, many of these methods have moderate attack success rates. Even when successful, the harmfulness of their outputs cannot be guaranteed, leading to suspicions that these methods have not accurately identified the safety vulnerabilities of LLMs. In this paper, we introduce a LLM attack method utilizing concept-based model explanation, where we extract safety concept activation vectors (SCAVs) from LLMs' activation space, enabling efficient attacks on well-aligned LLMs like LLaMA-2, achieving near 100% attack success rate as if LLMs are completely unaligned. This suggests that LLMs, even after thorough safety alignment, could still pose potential risks to society upon public release. To evaluate the harmfulness of outputs resulting with various attack methods, we propose a comprehensive evaluation method that reduces the potential inaccuracies of existing evaluations, and further validate that our method causes more harmful content. Additionally, we discover that the SCAVs show some transferability across different open-source LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xu et al. (2024) studied this question.

synapsesocial.com/papers/68e6e99bb6db643587664883https://doi.org/10.48550/arxiv.2404.12038
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming2024 · 3 citations
  2. 2Online Safety Analysis for LLMs: a Benchmark, an Assessment, and a Path Forward2024 · 1 citations
  3. 3Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation2024 · 1 citations
  4. 4LLM-Safety Evaluations Lack Robustness2025
  5. 5Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning2024 · 1 citations