The Pitfalls of KV Cache Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Alex, Geh, Renato, Grover, Aditya, Broeck, Guy Van den, Israel, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enabling Autoregressive Models to Fill In Masked Tokens
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
Adversarial Tokenization
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025)
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2024)
von: Zhao, Siyan, et al.
Veröffentlicht: (2024)
Collapsed Inference for Bayesian Deep Learning
von: Zeng, Zhe, et al.
Veröffentlicht: (2023)
von: Zeng, Zhe, et al.
Veröffentlicht: (2023)
On the Relationship Between Monotone and Squared Probabilistic Circuits
von: Wang, Benjie, et al.
Veröffentlicht: (2024)
von: Wang, Benjie, et al.
Veröffentlicht: (2024)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
How to Marginalize in Causal Structure Learning?
von: Zhao, William, et al.
Veröffentlicht: (2025)
von: Zhao, William, et al.
Veröffentlicht: (2025)
Scaling Up Probabilistic Circuits by Latent Variable Distillation
von: Liu, Anji, et al.
Veröffentlicht: (2022)
von: Liu, Anji, et al.
Veröffentlicht: (2022)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
Probabilistic Programs of Thought
von: Garg, Poorva, et al.
Veröffentlicht: (2026)
von: Garg, Poorva, et al.
Veröffentlicht: (2026)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
A Tractable Inference Perspective of Offline RL
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
von: Liu, Xuejie, et al.
Veröffentlicht: (2023)
Restructuring Tractable Probabilistic Circuits
von: Zhang, Honghua, et al.
Veröffentlicht: (2024)
von: Zhang, Honghua, et al.
Veröffentlicht: (2024)
SIMPLE: A Gradient Estimator for $k$-Subset Sampling
von: Ahmed, Kareem, et al.
Veröffentlicht: (2022)
von: Ahmed, Kareem, et al.
Veröffentlicht: (2022)
ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
A Pseudo-Semantic Loss for Autoregressive Models with Logical Constraints
von: Ahmed, Kareem, et al.
Veröffentlicht: (2023)
von: Ahmed, Kareem, et al.
Veröffentlicht: (2023)
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
Probabilistic Circuits for Cumulative Distribution Functions
von: Broadrick, Oliver, et al.
Veröffentlicht: (2024)
von: Broadrick, Oliver, et al.
Veröffentlicht: (2024)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
Where is the signal in tokenization space?
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
TurboAngle: Near-Lossless KV Cache Compression via Uniform Angle Quantization
von: Patel, Dipkumar
Veröffentlicht: (2026)
von: Patel, Dipkumar
Veröffentlicht: (2026)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
von: Wang, Xiao
Veröffentlicht: (2026)
von: Wang, Xiao
Veröffentlicht: (2026)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
von: Yang, Bin, et al.
Veröffentlicht: (2025)
von: Yang, Bin, et al.
Veröffentlicht: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enabling Autoregressive Models to Fill In Masked Tokens
von: Israel, Daniel, et al.
Veröffentlicht: (2025) -
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025) -
Adversarial Tokenization
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025) -
Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2024) -
Collapsed Inference for Bayesian Deep Learning
von: Zeng, Zhe, et al.
Veröffentlicht: (2023)