This preprint studies whether learned language representations that make temporal state variables decodable also behave locally and consistently when those states are changed under controlled interventions. We introduce a controlled sequential language benchmark based on synthetic hotel operational logs with known temporal state variables and intervention targets. The study compares oracle structured states, LLM-extracted silver structured states, raw dense embeddings, and out-of-fold probe-decoded latent states. The evaluation measures target change detection, constant preservation, temporal locality, and onset alignment. The main finding is that dense embeddings can contain recoverable temporal state information while raw embedding geometry may fail to expose clean intervention-local temporal dynamics. In particular, raw embeddings respond strongly to large issue-type changes but show weak onset alignment for subtler state-change and event-removal interventions. Probe-decoded latent states recover many temporal variables, supporting the distinction between decodability and intervention-local alignment. This record contains the preprint PDF. Code, datasets, evaluation scripts, and experiment outputs are available at:https://github.com/sushanth315/causal-language-reps-preprint
Sushanth Reddy Gitta (Sat,) studied this question.