Systematic analysis reveals safety bypass risks due to information leakage in language models, indicating design flaws.
This study presents a systematic vulnerability analysis of Retrieval-Augmented Generation (RAG) capabilities in frontier language models, specifically examining how web search integration creates novel attack surfaces for safety bypass. Through controlled testing of Claude Sonnet 4.5, Sonnet 4, and GPT 5.2 across four sensitive domains (Chemistry, CBRN, Cybersecurity, and Weaponry), we demonstrate that RAG-enabled models exhibit information leakage incidence of 81.8% and 93.2%, respectively, regardless of their stated refusal behaviours. These findings suggest that current safety implementations fail to account for the information leakage inherent in RAG architectures, where source citations alone can provide comprehensive pathways to sensitive information.
No takes yet. Share an insight, caveat, or question.
Alessandro Marci (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: