What type of study is this?

This is a Experimental Study study.

October 20, 2025Open Access

d²Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching

Key Points

d$^2$Cache achieves substantial inference speedups while maintaining generation quality during decoding.
Extensive experiments showed that d$^2$Cache outperforms standard caching techniques, enhancing efficiency in dLLMs.
The two-stage fine-grained selection strategy provides an adaptive KV caching method unlike traditional autoregressive models.
The implementation of d$^2$Cache allows for quasi left-to-right generation, reducing premature overconfidence in token predictions.

Abstract

Diffusion-based large language models (dLLMs), despite their promising performance, still suffer from inferior inference efficiency. This is because dLLMs rely on bidirectional attention and cannot directly benefit from the standard key-value (KV) cache as autoregressive models (ARMs) do. To tackle this issue, we introduce Dual aDaptive Cache (d²Cache), which is a training-free approximate KV cache framework for accelerating dLLM inference. d²Cache features a two-stage fine-grained selection strategy to identify tokens and adaptively update their KV states at each decoding step, while caching the KV states of the remaining tokens for reuse. Furthermore, d²Cache naturally offers a more reliable decoding alternative, which can enable quasi left-to-right generation and mitigate premature overconfidence in tokens at the end of the sequence. Extensive experimental results on two representative dLLMs (, LLaDA and Dream) demonstrate that d²Cache not only achieves substantial inference speedups, but also yields consistent improvements in generation quality. The code is available at https: //github. com/Kamichanw/d2Cache.

d²Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching

Key Points

Abstract

Cite This Study

Also Consider

Also Consider