Every large language model knows that "fire" co-occurs with "hot." None has ever perceived heat. We present the first system that bridges this gap by building a world small enough to be completely understood, rather than scaling models larger. A 228K-parameter Transformer is placed in a closed perceptual environment (103 entities, 13 sensory dimensions, 6 causal rules) and guided through a six-stage developmental pathway from perceptual completion to autonomous linguistic inference. Developmental order is causally necessary: removing the interaction phase eliminates causal reasoning entirely (accuracy 0.001); shuffling phase order collapses higher-order capacities while leaving lower-order ones intact. The minimal sufficient pathway (perception → naming → interaction) achieves higher causal reasoning than the full six-stage sequence, revealing capacity competition between cognitive breadth and depth. The model generalizes through per-dimension compositional rules rather than memorization (disambiguation ratio 5.0:1). Once grounded, language acquires autonomous operational capacity: with all perceptual input removed, causal reasoning remains at ceiling (1.000 ± 0.000 across five seeds). Every computational step is mechanistically traceable; attention specialization, information arrival, and causal flow are localized layer by layer. Across a 50-fold parameter range, the same capacities emerge: the bottleneck is developmental structure, not scale. No prior work has occupied this position: the intersection of full decomposability and the complete perception-to-language-to-autonomous-reasoning pathway.
Zhiwei Liu (Sun,) studied this question.