Randomized trial demonstrates effective threat detection in large language models, suggesting new security measures.
Large language models (LLMs) are increasingly deployed behind conversational and API-driven interfaces that expose them directly to untrusted natural-language input, making them susceptible to prompt injection, jailbreaking, and data-exfiltration attacks that conventional Web Application Firewalls (WAFs) are not designed to detect, since such attacks are semantically rather than syntactically anoma- lous. This paper presents a self-hosted security middleware that sits between client applications and a locally hosted LLaMA-3 model (served via Ollama) and performs multi-layer threat detection on every incoming prompt. The pipeline combines (i) leetspeak-aware keyword matching, (ii) Shannon-entropy based obfuscation detection, and (iii) LLM-based semantic classification, whose outputs are fused into a single weighted risk score that routes each request into one of three operating modes: SAFE, MONITOR, or DECEPTION. Unlike systems that merely block or warn, the DECEPTION mode actively engages suspected attackers with dynamically generated, plausible-looking fake credentials, database schemas, and system configuration data, functioning as a natural-language honeypot that wastes attacker effort while logging attacker behavior. The system further maintains a TTL-based semantic cache to reduce redundant analysis, a per-key sliding-window session store to detect multi-turn attacks, API-key authen- tication with rate limiting, an output-side data-loss-prevention (DLP) scanner that redacts personally identifiable information (PII) before responses leave the system, and a real-time monitoring dashboard. Every component is open source and runs entirely on local infrastructure, with no dependency on ex- ternal paid APIs. We describe the system’s architecture, scoring formula, and deception mechanism, and report on functional and adversarial testing performed against the implementation. We position this work as a practical, reproducible, and cost-free alternative to commercial LLM guardrail products, and discuss its current limitations and directions for quantitative evaluation.
No takes yet. Share an insight, caveat, or question.
A Nithin Kumar Shetty (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: