Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Yuerong, Liu, Xiaoran, Li, Ruixiao, Liu, Zhigeng, Huang, Zengfeng, Guo, Qipeng, He, Ziwei, Qiu, Xipeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915599637020672
author Song, Yuerong
Liu, Xiaoran
Li, Ruixiao
Liu, Zhigeng
Huang, Zengfeng
Guo, Qipeng
He, Ziwei
Qiu, Xipeng
author_facet Song, Yuerong
Liu, Xiaoran
Li, Ruixiao
Liu, Zhigeng
Huang, Zengfeng
Guo, Qipeng
He, Ziwei
Qiu, Xipeng
contents Diffusion Large Language Models (dLLMs) enable breakthroughs in reasoning and parallel decoding but suffer from prohibitive quadratic computational complexity and memory overhead during inference. Current caching techniques accelerate decoding by storing full-layer states, yet impose substantial memory usage that limit long-context applications. Our analysis of attention patterns in dLLMs reveals persistent cross-layer sparsity, with pivotal tokens remaining salient across decoding steps and low-relevance tokens staying unimportant, motivating selective cache eviction. We propose Sparse-dLLM, the first training-free framework integrating dynamic cache eviction with sparse attention via delayed bidirectional sparse caching. By leveraging the stability of token saliency over steps, it retains critical tokens and dynamically evicts unimportant prefix/suffix entries using an attention-guided strategy. Extensive experiments on LLaDA and Dream series demonstrate Sparse-dLLM achieves up to 10$\times$ higher throughput than vanilla dLLMs, with comparable performance and similar peak memory costs, outperforming previous methods in efficiency and effectiveness. The code is available at https://github.com/OpenMOSS/Sparse-dLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
Song, Yuerong
Liu, Xiaoran
Li, Ruixiao
Liu, Zhigeng
Huang, Zengfeng
Guo, Qipeng
He, Ziwei
Qiu, Xipeng
Computation and Language
Diffusion Large Language Models (dLLMs) enable breakthroughs in reasoning and parallel decoding but suffer from prohibitive quadratic computational complexity and memory overhead during inference. Current caching techniques accelerate decoding by storing full-layer states, yet impose substantial memory usage that limit long-context applications. Our analysis of attention patterns in dLLMs reveals persistent cross-layer sparsity, with pivotal tokens remaining salient across decoding steps and low-relevance tokens staying unimportant, motivating selective cache eviction. We propose Sparse-dLLM, the first training-free framework integrating dynamic cache eviction with sparse attention via delayed bidirectional sparse caching. By leveraging the stability of token saliency over steps, it retains critical tokens and dynamically evicts unimportant prefix/suffix entries using an attention-guided strategy. Extensive experiments on LLaDA and Dream series demonstrate Sparse-dLLM achieves up to 10$\times$ higher throughput than vanilla dLLMs, with comparable performance and similar peak memory costs, outperforming previous methods in efficiency and effectiveness. The code is available at https://github.com/OpenMOSS/Sparse-dLLM.
title Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
topic Computation and Language
url https://arxiv.org/abs/2508.02558