Randomized trial demonstrates efficient KV cache compression in large language models, highlighting improved memory usage.
The paper presents a hierarchical clustering framework for online Key-Value (KV) cache compression in Transformer-based large language models. The proposed approach organizes structurally similar KV representations into compact prototype clusters to reduce memory consumption while preserving contextual information during long-context autoregressive inference. The manuscript introduces a localized similarity-driven clustering strategy, Progressive Similarity Association (PSA), together with Adaptive Regional Partitioning (ARP) and Prototype Aggregation (PAG) to enable efficient online cache compression with low computational overhead. The paper also includes theoretical analysis, empirical observations, experimental evaluation, and ablation studies across long-context language modeling benchmarks. This deposit serves as an open-access research manuscript intended to facilitate discussion, reproducibility, and future research in efficient Transformer inference and long-context large language models.
No takes yet. Share an insight, caveat, or question.
Mukesh Anand G (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: