We argue that for a large language model (LLM) the externally writable in-context token sequence IS the cognitive substrate for category-use, not a representation OF cognition that runs on some deeper substrate, but the substrate itself, the only handle that is movable from outside the weights. We call this position the Token-Substrate Hypothesis (TSH). The strong form of Sapir-Whorf was rejected for humans because humans have prelinguistic cognition, the Off-Token Route. LLMs do not, and for them Wittgenstein's Tractatus 5. 6 stops being metaphor and becomes architecture. We test TSH with a methodology we call the Coinage Probe: a paired-trial elicitation that scores an LLM's distinguishability on a coined term against named near-neighbors before and after introducing a one-sentence canonical definition. Across 3 cross-vendor frontier models (Claude Opus 4. 7, GPT-5. 5, Gemini 2. 5 Pro) and 10 low-attestation coined targets plus 2 positive controls, we ran 108 trials x 3 near-neighbors per trial = 324 paired distinguishability measurements across 36 model x term cells, scored by a three-judge panel and compared against an author-rated 22-trial audit sample (panel-vs-author Cohen's κ = +0. 71). Mean cell-level Lexical Reachability (post minus cold) was +5. 47 on a 9-point scale (95% CI: +5. 13, +5. 80; cell-level Cohen's dcell = +3. 95 across n = 30 novel model x term cells). The effect replicated across the panel (cross-model CV = 0. 109) with a model-style interaction qualifying strict invariance, did not persist into a re-cold chat (H3 supported), and was absent on positive controls (H4 supported). These results support a bounded version of the Token-Substrate Hypothesis: in-context vocabulary functions as an externally writable substrate for LLM category use. For deployed LLM systems, notation is therefore not mere packaging; it is a design surface that shapes what distinctions the system can reliably use.
Alexandru Mares (Wed,) studied this question.