Training-time channel interventions are often motivated by dormant or apparently underused units, but this motivation can be misleading in small Transformer language models. This preprint studies MLP-channel gates in causal language models trained on WikiText-2, TinyStories, and TinyShakespeare. The results do not support a dormant-channel rejuvenation claim: targeted interventions do not consistently beat matched-random controls, and fixed low-duty-cycle gates can look useful under final-checkpoint selection while damaging the validation-loss trajectory. The supported claim is narrower: low-amplitude, strictly budgeted MLP channel gating reduces trajectory damage from fixed event gating and dropout in the tested small-model settings while keeping dense inference unchanged. The archive includes the LaTeX preprint, bibliography, generated tables and figures, diagnostics, scripts, run metadata, and reproducibility manifest.https://huggingface.co/datasets/cuber12/budgeted-guardrails-mlp-channel-interventions/tree/preprint-v2026-06-17 References Qiao, S., Lin, Z., Zhang, J., & Yuille, A. (2018). Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization. arXiv:1812.00481. https://doi.org/10.48550/arXiv.1812.00481 Sokar, G., Agarwal, R., Castro, P. S., & Evci, U. (2023). The Dormant Neuron Phenomenon in Deep Reinforcement Learning. arXiv:2302.12902. https://doi.org/10.48550/arXiv.2302.12902 Li, Z., You, C., Bhojanapalli, S., Li, D., Rawat, A. S., Reddi, S. J., Ye, K., Chern, F., Yu, F., Guo, R., & Kumar, S. (2022). The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers. arXiv:2210.06313. https://doi.org/10.48550/arXiv.2210.06313 Zhang, Z., Lin, Y., Liu, Z., Li, P., Sun, M., & Zhou, J. (2021). MoEfication: Transformer Feed-forward Layers are Mixtures of Experts. arXiv:2110.01786. https://doi.org/10.48550/arXiv.2110.01786 Zhang, Z., Xiao, C., Qin, Q., Lin, Y., Zeng, Z., Han, X., Liu, Z., Xie, R., Sun, M., & Zhou, J. (2024). Exploring the Benefit of Activation Sparsity in Pre-training. arXiv:2410.03440. https://doi.org/10.48550/arXiv.2410.03440 Dietrich, A., Gressmann, F., Orr, D., Chelombiev, I., Justus, D., & Luschi, C. (2021). Towards Structured Dynamic Sparse Pre-Training of BERT. arXiv:2108.06277. https://doi.org/10.48550/arXiv.2108.06277 Hu, P., Li, S., & Huang, L. (2024). Mixed Sparsity Training: Achieving 4x FLOP Reduction for Transformer Pretraining. arXiv:2408.11746. https://doi.org/10.48550/arXiv.2408.11746 Fan, A., Grave, E., & Joulin, A. (2019). Reducing Transformer Depth on Demand with Structured Dropout. arXiv:1909.11556. https://doi.org/10.48550/arXiv.1909.11556 Lin, Z., Liu, P., Huang, L., Chen, J., Qiu, X., & Huang, X. (2019). DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks. arXiv:1907.11065. https://doi.org/10.48550/arXiv.1907.11065 Zhou, W., Ge, T., Xu, K., Wei, F., & Zhou, M. (2020). Scheduled DropHead: A Regularization Method for Transformer Models. arXiv:2004.13342. https://doi.org/10.48550/arXiv.2004.13342 Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems. https://arxiv.org/abs/1706.03762 Geva, M., Schuster, R., Berant, J., & Levy, O. (2021). Transformer Feed-Forward Layers Are Key-Value Memories. Proceedings of EMNLP 2021. https://doi.org/10.18653/v1/2021.emnlp-main.446 Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A Simple Way to Prevent Neural Networks from Overfitting. Journal of Machine Learning Research, 15(56), 1929–1958. https://jmlr.org/papers/v15/srivastava14a.html Merity, S., Xiong, C., Bradbury, J., & Socher, R. (2016). Pointer Sentinel Mixture Models. arXiv:1609.07843. https://doi.org/10.48550/arXiv.1609.07843 Salesforce AI Research. (2026). WikiText Dataset Card. https://huggingface.co/datasets/Salesforce/wikitext Eldan, R., & Li, Y. (2023). TinyStories: How Small Can Language Models Be and Still Speak Coherent English? arXiv:2305.07759. https://doi.org/10.48550/arXiv.2305.07759 Eldan, R., & Li, Y. (2026). TinyStories Dataset Card. https://huggingface.co/datasets/roneneldan/TinyStories Karpathy, A. (2015). Tiny Shakespeare Dataset from char-rnn. https://github.com/karpathy/char-rnn/tree/master/data/tinyshakespeare Nair, V., & Hinton, G. E. (2010). Rectified Linear Units Improve Restricted Boltzmann Machines. Proceedings of ICML 2010. https://icml.cc/Conferences/2010/papers/432.pdf Hendrycks, D., & Gimpel, K. (2016). Gaussian Error Linear Units (GELUs). arXiv:1606.08415. https://doi.org/10.48550/arXiv.1606.08415 So, D. R., Mańke, W., Liu, H., Dai, Z., Shazeer, N., & Le, Q. V. (2021). Primer: Searching for Efficient Transformers for Language Modeling. arXiv:2109.08668. https://doi.org/10.48550/arXiv.2109.08668 Kingma, D. P., & Ba, J. (2014). Adam: A Method for Stochastic Optimization. arXiv:1412.6980. https://doi.org/10.48550/arXiv.1412.6980 Loshchilov, I., & Hutter, F. (2017). Decoupled Weight Decay Regularization. arXiv:1711.05101. https://doi.org/10.48550/arXiv.1711.05101 Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., & Chintala, S. (2019). PyTorch: An Imperative Style, High-Performance Deep Learning Library. Advances in Neural Information Processing Systems. https://papers.nips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library Apple Developer. (2026). Accelerated PyTorch Training on Mac. https://developer.apple.com/metal/pytorch/ Efron, B. (1979). Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552 Holm, S. (1979). A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733 Hollander, M., Wolfe, D. A., & Chicken, E. (2013). Nonparametric Statistical Methods (3rd ed.). Wiley.
Mandeep Sidhu (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: