CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Ning, Wang, Chengzhi, Liu, Yibo, Tian, Baoliang, Zhang, Haijun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
por: Yan, Xianglong, et al.
Publicado: (2025)
por: Yan, Xianglong, et al.
Publicado: (2025)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
por: Lin, Bokai, et al.
Publicado: (2024)
por: Lin, Bokai, et al.
Publicado: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025)
por: Tian, Yuxuan, et al.
Publicado: (2025)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
por: Zhang, Te, et al.
Publicado: (2025)
por: Zhang, Te, et al.
Publicado: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
por: Liu, Guangda, et al.
Publicado: (2024)
por: Liu, Guangda, et al.
Publicado: (2024)
The Pitfalls of KV Cache Compression
por: Chen, Alex, et al.
Publicado: (2025)
por: Chen, Alex, et al.
Publicado: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
por: Sun, Qiheng, et al.
Publicado: (2025)
por: Sun, Qiheng, et al.
Publicado: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
por: Xin, Jihao, et al.
Publicado: (2026)
por: Xin, Jihao, et al.
Publicado: (2026)
Palu: Compressing KV-Cache with Low-Rank Projection
por: Chang, Chi-Chih, et al.
Publicado: (2024)
por: Chang, Chi-Chih, et al.
Publicado: (2024)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
por: Zhao, Yi, et al.
Publicado: (2025)
por: Zhao, Yi, et al.
Publicado: (2025)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
por: Datta, Debajyoti, et al.
Publicado: (2026)
por: Datta, Debajyoti, et al.
Publicado: (2026)
Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs
por: Li, Guchan, et al.
Publicado: (2026)
por: Li, Guchan, et al.
Publicado: (2026)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
por: Zhou, Yuhao, et al.
Publicado: (2025)
por: Zhou, Yuhao, et al.
Publicado: (2025)
PatternKV: Flattening KV Representation Expands Quantization Headroom
por: Zhang, Ji, et al.
Publicado: (2025)
por: Zhang, Ji, et al.
Publicado: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
por: Wu, Wenbo, et al.
Publicado: (2025)
por: Wu, Wenbo, et al.
Publicado: (2025)
KVSculpt: KV Cache Compression as Distillation
por: Jiang, Bo, et al.
Publicado: (2026)
por: Jiang, Bo, et al.
Publicado: (2026)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
por: Swain, Kabir, et al.
Publicado: (2026)
por: Swain, Kabir, et al.
Publicado: (2026)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
por: Dong, Harry, et al.
Publicado: (2024)
por: Dong, Harry, et al.
Publicado: (2024)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
por: Chen, Chuangtao, et al.
Publicado: (2026)
por: Chen, Chuangtao, et al.
Publicado: (2026)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
por: Lu, Liming, et al.
Publicado: (2026)
por: Lu, Liming, et al.
Publicado: (2026)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
por: Li, Kunxi, et al.
Publicado: (2025)
por: Li, Kunxi, et al.
Publicado: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
por: Yang, Dongquan, et al.
Publicado: (2025)
por: Yang, Dongquan, et al.
Publicado: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
CASK: Core-Aware Selective KV Compression for Reasoning Traces
por: Kim, Buseong, et al.
Publicado: (2026)
por: Kim, Buseong, et al.
Publicado: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
por: Yang, June Yong, et al.
Publicado: (2024)
por: Yang, June Yong, et al.
Publicado: (2024)
Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement Learning
por: Pan, Haolin, et al.
Publicado: (2025)
por: Pan, Haolin, et al.
Publicado: (2025)
TurboAngle: Near-Lossless KV Cache Compression via Uniform Angle Quantization
por: Patel, Dipkumar
Publicado: (2026)
por: Patel, Dipkumar
Publicado: (2026)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
por: Ma, Xindian, et al.
Publicado: (2026)
por: Ma, Xindian, et al.
Publicado: (2026)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
por: Liu, Sihao, et al.
Publicado: (2026)
por: Liu, Sihao, et al.
Publicado: (2026)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
por: Yang, Yaoxin, et al.
Publicado: (2025)
por: Yang, Yaoxin, et al.
Publicado: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
por: Feng, Shaoting, et al.
Publicado: (2025)
por: Feng, Shaoting, et al.
Publicado: (2025)
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
por: Zhu, Rui, et al.
Publicado: (2026)
por: Zhu, Rui, et al.
Publicado: (2026)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
por: Wang, Zihan, et al.
Publicado: (2026)
por: Wang, Zihan, et al.
Publicado: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
por: Ahn, Jinwoo, et al.
Publicado: (2026)
por: Ahn, Jinwoo, et al.
Publicado: (2026)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
por: Zhang, Rongzhi, et al.
Publicado: (2024)
por: Zhang, Rongzhi, et al.
Publicado: (2024)
Ejemplares similares
-
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
por: Yan, Xianglong, et al.
Publicado: (2025) -
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
por: Lin, Bokai, et al.
Publicado: (2024) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025) -
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
por: Wang, Yixuan, et al.
Publicado: (2025) -
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
por: Zhang, Te, et al.
Publicado: (2025)