Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jianuo, Zhang, Yaojie, Yang, Yicun, Huang, Benhao, Qi, Biqing, Liu, Dongrui, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
by: Liu, Zhiyuan, et al.
Published: (2025)
by: Liu, Zhiyuan, et al.
Published: (2025)
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
by: Song, Yuerong, et al.
Published: (2025)
by: Song, Yuerong, et al.
Published: (2025)
FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration
by: Zhang, Yaojie, et al.
Published: (2026)
by: Zhang, Yaojie, et al.
Published: (2026)
Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
by: Jin, Xiangqi, et al.
Published: (2025)
by: Jin, Xiangqi, et al.
Published: (2025)
Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way
by: Yang, Yicun, et al.
Published: (2025)
by: Yang, Yicun, et al.
Published: (2025)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
by: Huang, Jianuo, et al.
Published: (2026)
by: Huang, Jianuo, et al.
Published: (2026)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
by: Long, Lingkun, et al.
Published: (2026)
by: Long, Lingkun, et al.
Published: (2026)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
by: Wei, Qingyan, et al.
Published: (2025)
by: Wei, Qingyan, et al.
Published: (2025)
Fast-dLLM v2: Efficient Block-Diffusion LLM
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
by: Feng, Yuan, et al.
Published: (2025)
by: Feng, Yuan, et al.
Published: (2025)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
by: Zhou, Enshuai, et al.
Published: (2026)
by: Zhou, Enshuai, et al.
Published: (2026)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
by: Li, Xuelin, et al.
Published: (2025)
by: Li, Xuelin, et al.
Published: (2025)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
by: Feng, Yuan, et al.
Published: (2024)
by: Feng, Yuan, et al.
Published: (2024)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026)
by: Yang, Xintong, et al.
Published: (2026)
dLLM: Simple Diffusion Language Modeling
by: Zhou, Zhanhui, et al.
Published: (2026)
by: Zhou, Zhanhui, et al.
Published: (2026)
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
by: Gu, Yifeng, et al.
Published: (2025)
by: Gu, Yifeng, et al.
Published: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
In-context KV-Cache Eviction for LLMs via Attention-Gate
by: Zeng, Zihao, et al.
Published: (2024)
by: Zeng, Zihao, et al.
Published: (2024)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
by: Liao, Mengqi, et al.
Published: (2025)
by: Liao, Mengqi, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
Elastic-dLLM: Position Preserving Context Compression and Augmentation of Diffusion LLMs
by: Wu, Junyi, et al.
Published: (2026)
by: Wu, Junyi, et al.
Published: (2026)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
by: Xiao, Zhongyu, et al.
Published: (2026)
by: Xiao, Zhongyu, et al.
Published: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
by: Chen, Jinhan, et al.
Published: (2025)
by: Chen, Jinhan, et al.
Published: (2025)
Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Similar Items
-
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
by: Liu, Zhiyuan, et al.
Published: (2025) -
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
by: Song, Yuerong, et al.
Published: (2025) -
FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration
by: Zhang, Yaojie, et al.
Published: (2026) -
Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
by: Jin, Xiangqi, et al.
Published: (2025) -
Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way
by: Yang, Yicun, et al.
Published: (2025)