Computational study demonstrates calibrated subspace gating prevents repetitive mode collapse in transformers, indicating geometric awareness stabilizes activation steering.
Key Points
To evaluate whether affinity-aware steering using calibrated Grassmannian overlap can prevent representational shear and degenerative repetitive mode collapse caused by static activation addition in transformer models.
Intervened at the Layer-14 residual stream of Qwen2.5-1.5B (hidden dimension d = 1536, 28 layers) using a difference-of-means steering vector (Euclidean norm = 21.0469) extracted from five contrastive pairs.
Constructed a rank-3 target attractor subspace from centered positive activations and a rank-2 state subspace using a sliding window delay embedding of width W = 4.
Calibrated raw Grassmannian overlap against ambient dimensionality using a temperature-scaled sigmoidal transfer function to expand the dynamic range.
Raw Grassmannian overlap concentrated near 0.008, matching the theoretical random-subspace floor (expected value ≈ 0.0039) and rendering uncalibrated gating ineffective.
Temperature-scaled sigmoidal calibration successfully restored the usable dynamic range to [0.41, 0.61], with a mean affinity of 0.526 on the test prompt.
Standard activation addition at coefficient c = 1.8 triggered severe phrase-level repetition under greedy decoding, whereas calibrated steering completely eliminated the loop by attenuating injection during syntactic transitions (affinity ≈ 0.41) and amplifying it during semantic predicates (affinity ≈ 0.61).