Randomized trial demonstrates improved performance in language model quantization, suggesting better efficiency for on-device applications.
Version 2 — supersedes v1. This version replaces the entropy-based method of v1 with a fundamentally different approach. Controlled experiments (Section 6) showed that v1's central claim does not hold: Shannon-entropy sensitivity ranking is statistically indistinguishable from random ranking on both tested models, and v1's reported gains are explained by a SmoothQuant confound (the baseline lacked the same pre-processing as the treatment arm). We publish this correction openly, with the controls that produced it. SGSR-2 replaces the heuristic proxy with direct measurement. For each transformer block and each configuration (bits ∈ {3,4,5,6} × group size ∈ {32,64,128}), we measure the KL divergence of output logits under reversible fake-quantization of that block alone. A validated additivity assumption (prediction/measurement ratio 0.74–1.11, rank order preserved) reduces budget-constrained allocation to an exact per-block Lagrangian rule, yielding the complete 3–5 bit/w Pareto frontier from a single overnight profiling run on consumer hardware (Apple M1 Pro, 32 GB). The method is gradient-free and runs entirely on-device. Under unified accounting (on-disk size over total parameters, full Wikitext-2 sliding-window perplexity with bootstrap confidence intervals), SGSR-2 strictly dominates uniform MLX quantization on both TinyLlama-1.1B and Qwen2.5-7B — e.g. +18.7% PPL at 3.62 bit/w versus +36.0% at 3.93 bit/w for the best uniform setting on Qwen — and matches or outperforms the hand-tuned llama.cpp K-quants frontier at ≥4.4 bit/w, while losing to Q3_K_M below 4 bit/w, a gap attributable to storage-format richness rather than allocation quality. Code, cost tables, all experimental results, and the evaluation harness are available at: https://github.com/Matth21/atlas
No takes yet. Share an insight, caveat, or question.
Matthias nicola Matthias Nicola Raviotta (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: