Attention Is All You Need for KV Cache in Diffusion LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen-Tri, Quan, Ranjan, Mukul, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
von: Gan, Chunjing, et al.
Veröffentlicht: (2024)
von: Gan, Chunjing, et al.
Veröffentlicht: (2024)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
von: Ghadia, Ravi, et al.
Veröffentlicht: (2025)
von: Ghadia, Ravi, et al.
Veröffentlicht: (2025)
Synthetic Data RL: Task Definition Is All You Need
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
von: Poudel, Pratik
Veröffentlicht: (2025)
von: Poudel, Pratik
Veröffentlicht: (2025)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
KV Cache Offloading for Context-Intensive Tasks
von: Bocharnikov, Andrey, et al.
Veröffentlicht: (2026)
von: Bocharnikov, Andrey, et al.
Veröffentlicht: (2026)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
von: He, Shenghua, et al.
Veröffentlicht: (2025)
von: He, Shenghua, et al.
Veröffentlicht: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
von: Guo, Zhiyu, et al.
Veröffentlicht: (2024)
von: Guo, Zhiyu, et al.
Veröffentlicht: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
KV Cache Transform Coding for Compact Storage in LLM Inference
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
von: Steinmetz, Cody, et al.
Veröffentlicht: (2025)
von: Steinmetz, Cody, et al.
Veröffentlicht: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024) -
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025) -
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026) -
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025) -
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)