NFV v1.2 introduces a trajectory-aware governance layer for conversational AI that augments static moderation with persistent risk modeling, adversarial escalation detection, and calibration-aware control logic. Unlike message-level classifiers that evaluate prompts in isolation, NFV models conversational trajectories across turns, tracking frame transitions, persistence of adversarial intent, and slow-burn escalation patterns. The framework formalizes a taxonomy of conversational frames, canonical trajectory risk tables, dual-window memory with exponential decay, and configurable governance profiles that balance safety and creative latitude. NFV is designed as a deployable governance middleware rather than a purely theoretical proposal. The specification includes explicit equations, reference pseudocode, audit reason codes, telemetry fields, and an evaluation protocol suitable for benchmarking against documented jailbreak datasets and benign conversational corpora. Key features include: Trajectory-aware adversarial risk scoring across conversational turns Persistent suspicion modeling with controlled decay and reset rules Canonical 2-gram and 3-gram trajectory risk tables Semantic similarity gating for paraphrase-resistant roleplay and boundary-testing detection Configurable governance profiles (STRICT, BALANCED, CREATIVE) Explicit failure modes, limitations, and deployment considerations This document consolidates NFV v1.2, v1.2.1, and v1.2.2 into a single unified and citable specification. It is intended for researchers, safety engineers, and system designers seeking robust, explainable, and production-viable conversational AI governance mechanisms.
Kon Lionis (2026) studied this question.