Language model output probability distributions directly reveal whether trained pathways exist in the model's weight matrices. At prompt tokens corresponding to known concepts, the model's predictions are near-deterministic; at fabricated concepts, predictions scatter across phonetic alternatives (d = 0.84, p 0.5), content authentication (d = 0.93), and manufacturing quality verification (d = 0.90) — all on a general-purpose model with zero domain-specific training. Additional applications include reasoning chain segmentation, through-thickness geometric profiling, training data composition detection, and a continuous memorization spectrum for IP forensics. The pathway specificity mechanism is argued to be structurally tamper-resistant: confident predictions at unfamiliar tokens require trained pathways, making score improvement indistinguishable from model improvement. Injection experiments support this argument, with architecture-dependent partial geometric shifts documented as a boundary condition. These findings support fourteen distinct applications and establish that external geometric measurement and internal mechanistic interpretability are complementary approaches to AI verification.
Joseph Stephens (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: