Abstract Background Large language models have developed a remarkable and largely misunderstood capacity: when confronted with a structured problem — whether algebraic, logical, or argumentative — they tend to produce outputs that are not merely fluent but semantically organized, reduced, and normalized. This phenomenon, which we term implicit predicate canonicalization, is not the product of explicitly programmed symbolic rules. It emerges spontaneously from the statistical regularities encoded during large-scale probabilistic training. Despite its pervasiveness, it remains one of the most poorly understood emergent properties of generative AI: it operates without any mechanism of control, without any guarantee of correctness, and without any formal grounding in the mathematical theory of canonical forms. The deep source of this limitation is architectural. Language generation systems reason at the level of the token — a sub-word fragment carrying zero intrinsic semantic content — and optimize for statistical plausibility, not propositional truth. A system that maximizes likelihood over structured texts learns to produce outputs that resemble canonical forms; it does not decide whether they are canonical forms. This distinction — between resembling a canonical form and being a canonical form — is the distinction between statistical discovery and formal decision. Hallucinations are not accidental errors: they are unstable linguistic trajectories, structurally inherent to a system deprived of any intrinsic logical convergence mechanism. The present article departs from a fundamental observation established across the S-AI corpus: the token is not the natural unit of reasoning. The predicate is. Three independent intellectual traditions — formal logic since Frege, universal grammar since Chomsky, and the Language of Thought since Fodor — converge toward the same conclusion: what reasoning operates on is not a word but a typed relation between arguments. The surface variability that characterizes natural language is precisely the lexical and syntactic noise enveloping an invariant predicate-argument structure. This noise is what current architectures treat as signal. This signal is what S-AI-PTR treats as primitive. Methods This article introduces S-AI-PTR (Predicate TRansformer), the ninth architectural instantiation of the Sparse Artificial Intelligence paradigm. It reconceptualizes the fundamental unit of language reasoning from the token to the canonical predicate. S-AI-PTR operates through a four-layer pipeline: a linguistic layer exploiting the LLM’s implicit canonicalization tendency to generate initial predicate hypotheses; a predicate extraction layer canonicalizing them into typed predicate tokens; a predicate reasoning and certification layer implementing the Predicate Reasoning Cycle with hormonal regulation and formal decidability; and a lexicalization layer producing the final text from the certified predicate skeleton. The architecture introduces two hormones specific to predicate reasoning — Predicatine, the predicate convergence hormone, and Ambiguine, the predicate uncertainty hormone — whose antagonistic coupling governs the PRC under the formal deployability condition. The full hormonal system comprises seven hormones whose coupled dynamics are analyzed by Lyapunov’s direct method. The resulting predicate skeleton constitutes the formal reasoning output: a certified Predicate Dependency Graph whose propositional content is grounded, coherent, and formally validated against the domain axiom base. The lexicalization layer dresses this skeleton in natural language, producing text that is formally constrained to express only what the skeleton contains. Results The theoretical framework establishes five principal formal results. Theorem PTR.1 proves the global asymptotic stability of the predicate hormonal subsystem under an analytically verifiable deployability condition, guaranteeing convergence from any initial state. Theorem PTR.2, the Predicate Entropic Contraction Theorem, establishes the formal equivalence between hormonal convergence and monotone predicate entropy reduction — the first explicitly linguistic instantiation of the governing doctrinal invariant of the S-AI paradigm. Theorem PTR.3 provides a finite-time termination guarantee with all three termination components analytically computable from architectural parameters prior to deployment. Theorem PTR.4 establishes the Predicate Decidability-Convergence Equivalence, bridging dynamical systems theory and computability theory in the predicate dimension. Theorem PTR.5 proves the global asymptotic stability of the complete hybrid LLM-PTR-Lexicalizer system. The evaluation framework formally defines seven metrics — Predicate Accuracy, Predicate Decidability Rate, Predicate Entropy Reduction, Predicate-Text Fidelity, Predicate Hallucination Rate, Predicate Frugality Index, and Lyapunov-Entropy Pearson Correlation — and derives analytically, from the five theorems, the target values each metric must achieve under the deployability condition. The experimental program, whose execution requires as a prerequisite the construction of a universal canonical predicate corpus, is formally specified with four benchmarks, seven reference systems, and a four-level ablation protocol. Conclusions The proposed framework defines a new class of language reasoning systems — Predicate Canonical Reasoning Systems (PCRS) — simultaneously satisfying five formally certified properties: predicate canonicality, logical validity, exponential convergence, predicate parsimony with fewer than ten million parameters, and guaranteed finite-time termination. These five properties are not independent engineering achievements: they are five expressions of a single underlying thermodynamic invariant — the Quintuple Equivalence of Predicate Reasoning — that unifies the dynamical, informational, symbolic, computational, and structural dimensions of predicate reasoning convergence under a single analytically verifiable condition. The central conceptual shift of the article is from generative models that discover canonical forms through statistical approximation to cognitive systems that decide whether those forms are formally correct through guaranteed dynamic convergence.
Said Slaoui (Sun,) studied this question.