We introduce TTAS-X2, a novel KV cache compression method that treats Keys (K) and Values (V) with fundamentally different strategies based on their functional roles in attention. While previous approaches compress K and V uniformly, we show that: K is directionâsensitive and must be preserved with high fidelity V tolerates aggressive compression without harming attention scores By applying 4âbit scalar quantization to K and hierarchical product quantization (256 â 64) with Hadamard transform and fixedâbudget outliers to V, we achieve: Attention cosine similarity â„ 0. 8 across all layers of Qwen3-32B Compression ratio > 10Ă for KV cache BPW < 1. 6 while maintaining functional behavior This work enables running 32Bâclass models on devices with 8GB RAM, previously impossible. đ§ 1. Introduction KV cache is the main memory bottleneck during longâcontext inference. Standard quantization (2â3 bits) degrades attention scores severely because K is distorted, and softmax is highly sensitive to directional changes. We propose a functional separation: K: 4âbit scalar + perâhead scale V: RPQ + Hadamard + fixedâbudget outliers âïž 2. Methodology 2. 1 K Compression For each head, we store a 4âbit quantized version with a perâhead scale. No Hadamard, no PQ, no residuals â direction is preserved. 2. 2 V Compression RMS normalize per vector Hadamard transform to spread energy Twoâlevel product quantization (K1=256, K2=64â512) Fixedâbudget outliers (32â256 per head) đ 3. Experiments Model: Qwen3-32B (headdim = 128, 64 layers, GQA with 8 KV heads) Sequence length: 1024 tokens per test Layer Cosine OUTLIERBUDGET K2 0 0. 824 32 64 1 0. 804 64 128 2 0. 803 128 256 4 0. 806 256 512 All layers achieve cosine â„ 0. 8, proving that TTAS-X2 generalizes across depth. đŸ 4. Memory Analysis For a 32B model with 64 layers and GQA (8 KV heads): Context Length Original (FP16) TTAS-X2 Ratio BPW 4k 1. 07 GB 0. 17 GB 6. 3Ă 1. 59 16k 4. 29 GB 0. 68 GB 6. 3Ă 1. 59 32k 8. 59 GB 1. 36 GB 6. 3Ă 1. 59 64k 17. 18 GB 2. 72 GB 6. 3Ă 1. 59 TTAS-X2 enables 32k context on 8GB devices for the first time. đŹ 5. Conclusion We demonstrate that functional separation of K and V is the key to extreme KV cache compression. TTAS-X2 preserves attention fidelity while reducing memory footprint by an order of magnitude. Patent pending â all rights reserved. đ 6. References Multiverse Computing, "CompactifAI" (2025) Intel, "KV Cache Compression via Token Pruning" (2024) Original TTAS work (Alsheck, 2025) đŹ Contact Abdulaziz Alsheckđ§ abdulazizabdullahalalsheck@gmail. com Phone Number +966560756695
alsheck abdulaziz (2026) studied this question.