© Copyright 2026 Chen, Ho Yiing (ORCID: 0009-0006-6816-9891), Independent researcher, charenix. com. All rights reserved by the author. This work is licensed under the Creative Commons Attribution 4. 0. Two members of the qwen2. 5 family share an output-side grammar at the per-coordinate level. Aggregating activation heat at block L-1 for qwen2. 5: 7B (D₇=3584) and qwen2. 5: 3B (D₃=2048) on ten multilingual prompts via Mercury, I take each network's hottest hidden-state indices: 200 for the larger member, a proportionally scaled 114 for the smaller. The two top-K sets share 49 indices inside the smaller variant's range. Under a permutation null this is 4. 4x chance. Eleven of those 49 also host 7B's constant-firing set U, defined as cells with the all-ones signature across the prompt panel. A hypergeometric test against the full 3584-wide range gives p < 1. 2 x 10^-12 for this second overlap, 20. 6x chance. The eleven triple-confirmed slots are T = 11, 12, 25, 279, 334, 382, 476, 481, 510, 715, 758: addressable integer coordinates that are hot in 7B aggregate, hot in 3B aggregate (after proportional top-K scaling), and active on every one of the ten observation prompts when looked up in U. Practical readings: mergekit-style weight surgery between qwen2. 5 members should protect these eleven slots; distillation from 7B to 3B should weight alignment loss higher along these channels; SAE-style interpretability work has an a-priori target subset for monosemantic feature probes. Companion records: Mercury method (10. 5281/zenodo. 20313154), Mercury discovery (10. 5281/zenodo. 20313748), Mercury-Viewer dataset (10. 5281/zenodo. 20313150), Mercury scaling (10. 5281/zenodo. 20325676).
Ho Yiing Chen (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: