PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 20240 citationsOpen Access

Exploring Safety Generalization Challenges of Large Language Models via Code

View Full Paper
QRQibing RenCGChang GaoXidian UniversityJSJing ShaoHong Kong Baptist University

Key Points

Key points are not available for this paper at this time.

Abstract

The rapid advancement of Large Language Models (LLMs) has brought about remarkable capabilities in natural language processing but also raised concerns about their potential misuse. While strategies like supervised fine-tuning and reinforcement learning from human feedback have enhanced their safety, these methods primarily focus on natural languages, which may not generalize to other domains. This paper introduces CodeAttack, a framework that transforms natural language inputs into code inputs, presenting a novel environment for testing the safety generalization of LLMs. Our comprehensive studies on state-of-the-art LLMs including GPT-4, Claude-2, and Llama-2 series reveal a common safety vulnerability of these models against code input: CodeAttack consistently bypasses the safety guardrails of all models more than 80\% of the time. Furthermore, we find that a larger distribution gap between CodeAttack and natural language leads to weaker safety generalization, such as encoding natural language input with data structures or using less popular programming languages. These findings highlight new safety risks in the code domain and the need for more robust safety alignment algorithms to match the code capabilities of LLMs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ren et al. (2024) studied this question.

synapsesocial.com/papers/68e747e6b6db6435876c0d53https://doi.org/10.48550/arxiv.2403.07865
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming2024 · 3 citations
  2. 2Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMs2024 · 7 citations
  3. 3Large Language Models in Code Co-generation for Safe Autonomous Vehicles2025
  4. 4A Survey on Large Language Models in Software Security: Opportunities and Threats2026 · 2 citations
  5. 5Exploring the Adversarial Capabilities of Large Language Models2024 · 1 citations