Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miao, Ruijie, Wang, Zhiming, Li, Wang, Wu, Shiwei, Liu, Shufan, Jiang, Yanbing, Yang, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
von: Shen, Yiqun, et al.
Veröffentlicht: (2025)
von: Shen, Yiqun, et al.
Veröffentlicht: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
A Simple Plug-in for Improving Eviction-Based KV Cache Compression
von: Lin, Yuping, et al.
Veröffentlicht: (2026)
von: Lin, Yuping, et al.
Veröffentlicht: (2026)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
von: Akulov, Dmitry, et al.
Veröffentlicht: (2025)
von: Akulov, Dmitry, et al.
Veröffentlicht: (2025)
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
von: Bui, Ngoc, et al.
Veröffentlicht: (2026)
von: Bui, Ngoc, et al.
Veröffentlicht: (2026)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
In-context KV-Cache Eviction for LLMs via Attention-Gate
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
von: Miao, Ruijie, et al.
Veröffentlicht: (2025)
von: Miao, Ruijie, et al.
Veröffentlicht: (2025)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
von: Fang, Shaoke, et al.
Veröffentlicht: (2026)
von: Fang, Shaoke, et al.
Veröffentlicht: (2026)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
von: Yang, Jiaming, et al.
Veröffentlicht: (2026)
von: Yang, Jiaming, et al.
Veröffentlicht: (2026)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
von: Li, Kunjun, et al.
Veröffentlicht: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
von: Shen, Yiqun, et al.
Veröffentlicht: (2025) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025) -
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
von: Zhou, Enshuai, et al.
Veröffentlicht: (2026) -
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
von: Li, Yubo, et al.
Veröffentlicht: (2026) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)