KVSculpt: KV Cache Compression as Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Bo, Jin, Sian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
KV Cache Offloading for Context-Intensive Tasks
von: Bocharnikov, Andrey, et al.
Veröffentlicht: (2026)
von: Bocharnikov, Andrey, et al.
Veröffentlicht: (2026)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
von: Jegou, Simon, et al.
Veröffentlicht: (2026)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
KV Cache Transform Coding for Compact Storage in LLM Inference
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
von: Staniszewski, Konrad, et al.
Veröffentlicht: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
Training-Free Exponential Context Extension via Cascading KV Cache
von: Willette, Jeffrey, et al.
Veröffentlicht: (2024)
von: Willette, Jeffrey, et al.
Veröffentlicht: (2024)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
von: Cai, Zhihang, et al.
Veröffentlicht: (2025)
von: Cai, Zhihang, et al.
Veröffentlicht: (2025)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
von: Wang, Luning, et al.
Veröffentlicht: (2024)
von: Wang, Luning, et al.
Veröffentlicht: (2024)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
von: Ghadia, Ravi, et al.
Veröffentlicht: (2025)
von: Ghadia, Ravi, et al.
Veröffentlicht: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
von: Poudel, Pratik
Veröffentlicht: (2025)
von: Poudel, Pratik
Veröffentlicht: (2025)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025) -
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026) -
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024) -
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)