Transformer language models generate text through autoregressive next-token prediction, whereas hallucination is a semantic and evidence-relative failure that generally becomes meaningful only at the level of propositions or claims. This mismatch makes it difficult to treat hallucination as an ordinary token classification problem. Existing studies show that uncertainty measures, hidden activations, attention patterns, and intervention directions can contain useful signals about factuality, yet the literature does not establish a single universal internal variable that corresponds to hallucination. This paper proposes a falsifiable theoretical abstraction in which hallucination is modeled as a latent, token-conditioned hazard process. The proposed Latent Hallucination Hazard Model (LHHM) defines a local hazard ᵢ over generation steps, conditioned on predictive uncertainty, generation commitment, hidden-state trajectory, evidence support, contextual conflict, and optional factuality probes. Local hazards are aggregated into claim-level risk through a survival-style formulation. The model explains why token entropy or attention mass alone cannot be sufficient, accommodates both uncertain confabulation and confident falsehood, and supports an architectural extension called the Hallucination State Head. A training objective, inference-time control policy, testable hypotheses, falsification criteria, and empirical validation protocol are specified. The paper does not claim that current Transformers possess a unique hallucination neuron; rather, it offers a disciplined mathematical framework for testing whether hallucination risk can be represented, calibrated, and controlled during autoregressive generation.
Kishore Chalakkal Varghese (Sun,) studied this question.