Large Language Models increasingly operate in contexts where outputs do more than inform: they trigger actions, authorize decisions, and persist as binding state. Yet the boundary at which raw signals—text, pixels, events—are allowed to count remains largely ungoverned. Tokenization is treated as plumbing: a compression step optimized for throughput, not legitimacy. This paper introduces Governed Semantic Tokens (GST), a framework that reframes tokenization as an admissibility boundary—the jurisdiction where meaning is either lawfully admitted or refused. Rather than opaque, model-specific vocabularies, GST employs registry-backed token identifiers whose semantics are explicit, typed, versioned, and governed. Tokens become units of admissible meaning, not merely units of prediction. By relocating governance to the tokenizer layer, GST enables structural resistance to prompt injection, cross-model semantic stability, auditable promotion to action, and refusal as a first-class success condition. The framework is model-agnostic and applies uniformly across modalities, treating text, vision, UI events, and tool interactions as inputs to a single governed semantic substrate. The paper positions GST as MCIR-0: the minimal viable substrate beneath higher-order governed systems, where meaning itself must be admitted before reasoning or action can proceed. Tokenization, long assumed neutral, is shown to be constitutional infrastructure.
Adam Ableman Mazurk (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: