Randomized trial demonstrates causal testing and computation insights in pretrained transformers, suggesting improved interpretability.
Mechanistic interpretability needs representations that do more than identify important components: they should expose the computation and support causal testing inside the original model. We introduce executable matrix programs, a weight-derived representation of attention and MLP components that requires no retraining or learned surrogate. For RoPE attention, each head is decomposed into relative-position QK routing programs and an affine VO payload map. The QK score can be unpacked into constant, query-affine, key-affine, and bilinear token-interaction terms, while the VO map shows what transformed payload is written to the residual stream. A SwiGLU MLP is represented as a sum of gated rank-1 atoms with explicit gate, read, and write factors. These objects are reused across inputs and can be executed inside the native forward pass. On Qwen2.5-0.5B-Instruct, simultaneous replacement of all 336 attention heads and all 24 MLP sublayers preserves the top-1 prediction on 64/64 prompts with a median full-logit discrepancy of 0.258%, validating the executable representation. The same interface supports separate interventions on QK routing, VO payload, and MLP computation, as well as downstream propagation analysis. This makes it possible to inspect why a token is selected, what information is transferred, which MLP atoms contribute, and how these factors affect later layers and final logits. Code and machine-readable results:https://github.com/maxwelhelp/matrix-programs
No takes yet. Share an insight, caveat, or question.
Maxim Vladimirovich Zhivotok (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: