PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 21, 20240 citationsOpen Access

Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content

View Full Paper
FBFederico BianchiJZJames Zou

Key Points

Key points are not available for this paper at this time.

Abstract

The risks derived from large language models (LLMs) generating deceptive and damaging content have been the subject of considerable research, but even safe generations can lead to problematic downstream impacts. In our study, we shift the focus to how even safe text coming from LLMs can be easily turned into potentially dangerous content through Bait-and-Switch attacks. In such attacks, the user first prompts LLMs with safe questions and then employs a simple find-and-replace post-hoc technique to manipulate the outputs into harmful narratives. The alarming efficacy of this approach in generating toxic content highlights a significant challenge in developing reliable safety guardrails for LLMs. In particular, we stress that focusing on the safety of the verbatim LLM outputs is insufficient and that we also need to consider post-hoc transformations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bianchi et al. (2024) studied this question.

synapsesocial.com/papers/68e785a2b6db6435876f7fbchttps://doi.org/10.48550/arxiv.2402.13926
Ask AI
Helpful
Bookmark
Share
View Full Paper