Key points are not available for this paper at this time.
Runtime guardrails for LLM agents enforce safety per individual action at execution time, but fail to verify the reachable state space before deployment. We present BeVerA-KG, a behaviour-driven framework that shifts verification to design time. Stakeholder-authored Gherkin scenarios are deterministically compiled into OWL 2 DL disjointness axioms (HermiT-verified for schema consistency) and paired with SPARQL 1.1 ASK guards for runtime defence-in-depth. The composed agent model is statistically analyzed as a priced timed automaton via UPPAAL statistical model checking (≥26,500 Monte Carlo runs, providing ≥ 0.99 confidence with ≤ 0.01 error per the Chernoff-Hoeffding bound), and re-verified post-quantization to ensure safety invariants survive deployment on ultra-constrained humanitarian edge hardware (<150 MB footprint, <200 ms latency for the symbolic engine). We formalise the statistical assurance as Pr□ϕ(s)∧♦ψ(s) ≥ 0.99 over reachable states—providing broader pre-deployment assurance than per-action enforcement alone—and contribute a twelve-scenario guardrail suite grounded in field census data, a machine-auditable test harness, and an open-source compliance pipeline.
Behailu Wolde (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: