Orthographic Perturbations in Qwen: A Replicated SAE Case Study with a Tokenizer-Equivalence Audit Protocol. Preprint, version 1.0 (June 2026). Not peer reviewed. Human-readable orthographic perturbations can preserve apparent prompt intent while changing the token sequence a model actually receives. We examine that mismatch in Qwen3.5-35B-A3B, inspecting the model with Qwen-Scope residual sparse autoencoders. For each of four prompt families we start from a plain ASCII prompt and compare it against perturbed and control rewrites: we check how the model's native tokenizer splits each one, and how far each rewrite moves the residual stream and the set of active SAE features. Visually similar prompts turn into different token sequences, and the rewrites produce measurable displacement in both the residual stream and the SAE feature neighborhood. Controls that match the perturbations on token count, add unusual ASCII, substitute Unicode nonletters, or shuffle word order move the representation about as much as dense diacritics do, so diacritics are only one of several contributing factors. These measurements are stable across three repeated deterministic runs of all 48 prompts and reproduce an earlier run up to floating-point noise. We keep the claim Qwen-specific: in this model, visible-text equivalence is not sufficient to establish tokenizer or representation equivalence. We frame tokenizer-induced non-equivalence (TINE) as a prospective audit for other models. This deposit contains the paper (LaTeX source and built PDF). The full evidence package (prompts, tokenizer audits, residual and SAE metric tables, run manifests, logs, checksums, provenance records, and analysis/figure scripts) is in the companion repository: https://github.com/jeffreywilliamportfolio/orthographic-effects-qwen-35b-a3b-sae
Jeffrey W. Shorthill (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: