Randomized trial shows improved quantization efficiency in large language models, suggesting enhanced model performance.
We present SGSR (Sensitivity-Guided Group-Size Redistribution), a post-training quantization method that uses per-layer Shannon entropy to redistribute quantization group sizes across transformer layers, rather than varying bit widths. We further introduce SGSR-Q, which combines group-size redistribution with a quality gate that promotes the top 5% most sensitive layers to 8-bit precision. On TinyLlama 1.1B, SGSR achieves +3.28% PPL delta at 4.51 bit/w versus +4.34% for uniform 4-bit quantization. On Qwen2.5-7B-Instruct, SGSR-Q achieves +6.72% PPL delta at 4.14 bit/w versus +11.78% uniform — a 43% relative improvement consistent across model scales. Both algorithms are implemented natively in MLX for Apple Silicon with no gradient computation required.
No takes yet. Share an insight, caveat, or question.
Matthias nicola Matthias Nicola Raviotta (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: