This preprint presents Latent Safety Margins (LSM), a three-state framework for large language model safety. The central hypothesis is that the conventional binary distinction between safe and harmful representations may be too restrictive for ambiguous, dual-use, delayed-intent, and progressively adversarial interactions. LSM introduces an explicit Uncertain representational regime between Safe and Harmful states. This regime is intended to preserve unresolved intent, delay premature safety decisions, and allow evidence to accumulate during autoregressive generation. The proposed framework combines a tri-state representation objective, a geometric margin regularizer, temporal supervision, and an evidence-accumulation gate for runtime monitoring of hidden-state trajectories. The manuscript defines falsifiable hypotheses and a reproducible experimental protocol for evaluating jailbreak robustness, false-refusal rates, calibration, early-warning lead time, and representation-space separability. This version is a research proposal and experimental protocol.
Marina Khorkina (Mon,) studied this question.