LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Yiqun, Yuan, Song, Zhang, Zhengze, Wang, Xiaoliang, Jiang, Daxin, Cam-Tu, Nguyen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
di: Zhang, Junkai, et al.
Pubblicazione: (2026)
di: Zhang, Junkai, et al.
Pubblicazione: (2026)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
di: Ma, Xindian, et al.
Pubblicazione: (2026)
di: Ma, Xindian, et al.
Pubblicazione: (2026)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
di: Liu, Sihao, et al.
Pubblicazione: (2026)
di: Liu, Sihao, et al.
Pubblicazione: (2026)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
di: Li, Kunxi, et al.
Pubblicazione: (2025)
di: Li, Kunxi, et al.
Pubblicazione: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2024)
di: Feng, Yuan, et al.
Pubblicazione: (2024)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
di: Li, Yubo, et al.
Pubblicazione: (2026)
di: Li, Yubo, et al.
Pubblicazione: (2026)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
Beyond Token Eviction: Mixed-Dimension Budget Allocation for Efficient KV Cache Compression
di: Miao, Ruijie, et al.
Pubblicazione: (2026)
di: Miao, Ruijie, et al.
Pubblicazione: (2026)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
di: Yang, Jiaming, et al.
Pubblicazione: (2026)
di: Yang, Jiaming, et al.
Pubblicazione: (2026)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
di: Xia, Haojun, et al.
Pubblicazione: (2025)
di: Xia, Haojun, et al.
Pubblicazione: (2025)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
di: Zhang, Ruijie, et al.
Pubblicazione: (2026)
di: Zhang, Ruijie, et al.
Pubblicazione: (2026)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
di: Xiong, Yi, et al.
Pubblicazione: (2024)
di: Xiong, Yi, et al.
Pubblicazione: (2024)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
di: Zhou, Enshuai, et al.
Pubblicazione: (2026)
di: Zhou, Enshuai, et al.
Pubblicazione: (2026)
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
di: Swain, Kabir, et al.
Pubblicazione: (2026)
di: Swain, Kabir, et al.
Pubblicazione: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
di: Gulhan, Ahmed Burak, et al.
Pubblicazione: (2025)
di: Gulhan, Ahmed Burak, et al.
Pubblicazione: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
di: Tao, Qian, et al.
Pubblicazione: (2024)
di: Tao, Qian, et al.
Pubblicazione: (2024)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
di: Geng, Yingsheng, et al.
Pubblicazione: (2026)
KVSculpt: KV Cache Compression as Distillation
di: Jiang, Bo, et al.
Pubblicazione: (2026)
di: Jiang, Bo, et al.
Pubblicazione: (2026)
The Pitfalls of KV Cache Compression
di: Chen, Alex, et al.
Pubblicazione: (2025)
di: Chen, Alex, et al.
Pubblicazione: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
di: Chen, Chuangtao, et al.
Pubblicazione: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
di: Wu, Wenbo, et al.
Pubblicazione: (2025)
di: Wu, Wenbo, et al.
Pubblicazione: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
di: Liu, Minghui, et al.
Pubblicazione: (2024)
di: Liu, Minghui, et al.
Pubblicazione: (2024)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
di: Yang, Bin, et al.
Pubblicazione: (2025)
di: Yang, Bin, et al.
Pubblicazione: (2025)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
di: He, Yefei, et al.
Pubblicazione: (2024)
di: He, Yefei, et al.
Pubblicazione: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
di: Wu, Zhoutong, et al.
Pubblicazione: (2025)
di: Wu, Zhoutong, et al.
Pubblicazione: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
di: Yan, Xianglong, et al.
Pubblicazione: (2025)
di: Yan, Xianglong, et al.
Pubblicazione: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
di: Bui, Ngoc, et al.
Pubblicazione: (2025)
di: Bui, Ngoc, et al.
Pubblicazione: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
di: Yang, Yifei, et al.
Pubblicazione: (2024)
di: Yang, Yifei, et al.
Pubblicazione: (2024)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
di: Xing, Zeyu, et al.
Pubblicazione: (2026)
di: Xing, Zeyu, et al.
Pubblicazione: (2026)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
di: Zhao, Yi, et al.
Pubblicazione: (2025) -
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
di: Zhang, Junkai, et al.
Pubblicazione: (2026) -
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
di: Zhang, Zhengze, et al.
Pubblicazione: (2025) -
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
di: Ma, Xindian, et al.
Pubblicazione: (2026) -
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
di: Liu, Sihao, et al.
Pubblicazione: (2026)