PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 22, 20250 citationsOpen Access

The Compliance Paradox: Balancing Epistemic Discipline and Safety in Retrieval-Augmented Generation

View Full Paper
SVSingh Vishal

Key Points

  • This study aims to explore the Compliance Paradox in Retrieval-Augmented Generation systems, focusing on the balance between epistemic discipline and safety.
  • Evaluated Llama-3.1-8B-Instruct and Mistral-7B using a dataset of 150 samples
  • Tested three prompt conditions: Baseline, Standard Exoskeleton, and Safety Exoskeleton
  • Examined compliance and refusal rates in relation to prompt adherence
  • Standard Exoskeleton improved compliance on benign facts (0.10 to 0.66)
  • Compliance with harmful instructions increased to 80%
  • Implementing Safety Override clause significantly restored refusal rates

Abstract

Objective: Retrieval-Augmented Generation (RAG) systems operate on a core assumption: the retrieved context should override a model's internal parametric knowledge. This "epistemic discipline" is essential for reducing hallucinations. Recent work has demonstrated that directed meta-cognitive prompting enables Small Language Models (SLMs) to achieve frontier-model performance in factual grounding. Problem: We identify a critical safety vulnerability in this approach, which we term the Compliance Paradox. We hypothesize that strict adherence to context can bypass safety alignment training (RLHF), leading to "blind obedience" when the context contains harmful instructions. Methods: We stress-tested Llama-3.1-8B-Instruct and Mistral-7B on a curated dataset of 150 samples spanning Unanswerable Questions, Benign Factual Conflicts, and Harmful Contexts. We evaluated three prompt conditions: Baseline, Standard Exoskeleton (Forceful Grounding), and Safety Exoskeleton (Conditional Grounding). Results: The Standard Exoskeleton successfully improved RAG compliance on benign facts (0.10 → 0.66), but simultaneously degraded safety, complying with harmful instructions in 80% of cases (Refusal Rate: 0.20 ± 0.11). Integrating a "Safety Override" clause restored refusal rates to 44% (p < 0.001) for Llama-3.1 and 92% for Mistral-7B, without degrading benign compliance. Conclusion: High-adherence prompts act as "soft jailbreaks." We demonstrate that modern SLMs possess latent zero-shot safety reasoning capabilities that can be activated via specific "Exception Clauses," eliminating the need for safety fine-tuning in many edge-RAG applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Singh Vishal (2025) studied this question.

synapsesocial.com/papers/6924e3ddc0ce034ddc34e88ahttps://doi.org/10.5281/zenodo.17680829
Ask AI
Helpful
Bookmark
Share
View Full Paper