Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chari, Vivek, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
von: Devoto, Alessio, et al.
Veröffentlicht: (2025)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
CommVQ: Commutative Vector Quantization for KV Cache Compression
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
von: Yuan, Aomufei, et al.
Veröffentlicht: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
von: Li, Mengjie, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
ZSMerge: Zero-Shot KV Cache Compression for Memory-Efficient Long-Context LLMs
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
More Than a Quick Glance: Overcoming the Greedy Bias in KV-Cache Compression
von: Sood, Aryan, et al.
Veröffentlicht: (2026)
von: Sood, Aryan, et al.
Veröffentlicht: (2026)
A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
LM Agents for Coordinating Multi-User Information Gathering
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
Towards Threshold-Free KV Cache Pruning
von: Ni, Xuanfan, et al.
Veröffentlicht: (2025)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2025)
KV Cache Steering for Controlling Frozen LLMs
von: Belitsky, Max, et al.
Veröffentlicht: (2025)
von: Belitsky, Max, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression
von: Goldstein, Daniel, et al.
Veröffentlicht: (2024)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2024)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
von: Chari, Vivek, et al.
Veröffentlicht: (2025) -
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025) -
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
von: Devoto, Alessio, et al.
Veröffentlicht: (2025) -
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024) -
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)