WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Zuo, Youhui, Wei, Sibo, Zhang, Chen, Liu, Zhuorui, Lu, Wenpeng, Song, Dawei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Long-Context LLM Inference via KV Cache Clustering
por: Hu, Jie, et al.
Publicado: (2025)
por: Hu, Jie, et al.
Publicado: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
por: Tao, Wei, et al.
Publicado: (2026)
por: Tao, Wei, et al.
Publicado: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
por: Yang, Yifei, et al.
Publicado: (2024)
por: Yang, Yifei, et al.
Publicado: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression
por: Li, Runchao, et al.
Publicado: (2025)
por: Li, Runchao, et al.
Publicado: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025)
por: Tian, Yuxuan, et al.
Publicado: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024)
por: Zhou, Xiabin, et al.
Publicado: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
por: Yu, Bohan, et al.
Publicado: (2025)
por: Yu, Bohan, et al.
Publicado: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
por: Sun, Hanshi, et al.
Publicado: (2024)
por: Sun, Hanshi, et al.
Publicado: (2024)
QAQ: Quality Adaptive Quantization for LLM KV Cache
por: Dong, Shichen, et al.
Publicado: (2024)
por: Dong, Shichen, et al.
Publicado: (2024)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
por: He, Xingyang, et al.
Publicado: (2025)
por: He, Xingyang, et al.
Publicado: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
por: Li, Kunxi, et al.
Publicado: (2025)
por: Li, Kunxi, et al.
Publicado: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
por: Gu, Yifeng, et al.
Publicado: (2025)
por: Gu, Yifeng, et al.
Publicado: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
por: Guo, Jinyu, et al.
Publicado: (2026)
por: Guo, Jinyu, et al.
Publicado: (2026)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
por: Liu, Zhuorui, et al.
Publicado: (2025)
por: Liu, Zhuorui, et al.
Publicado: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
por: Feng, Yuan, et al.
Publicado: (2025)
por: Feng, Yuan, et al.
Publicado: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
por: Li, Xing, et al.
Publicado: (2025)
por: Li, Xing, et al.
Publicado: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
por: Lu, Liming, et al.
Publicado: (2026)
por: Lu, Liming, et al.
Publicado: (2026)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
por: Dehghanighobadi, Zahra, et al.
Publicado: (2026)
por: Dehghanighobadi, Zahra, et al.
Publicado: (2026)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
por: Sharma, Akshat, et al.
Publicado: (2024)
por: Sharma, Akshat, et al.
Publicado: (2024)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
por: Ma, Da, et al.
Publicado: (2024)
por: Ma, Da, et al.
Publicado: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
por: Kang, Hao, et al.
Publicado: (2024)
por: Kang, Hao, et al.
Publicado: (2024)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)
por: Ji, Shiyu, et al.
Publicado: (2026)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
por: Li, Xuelin, et al.
Publicado: (2025)
por: Li, Xuelin, et al.
Publicado: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
por: Behnam, Payman, et al.
Publicado: (2025)
por: Behnam, Payman, et al.
Publicado: (2025)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
por: Wan, Zhongwei, et al.
Publicado: (2025)
por: Wan, Zhongwei, et al.
Publicado: (2025)
EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
por: Guo, Tianyu, et al.
Publicado: (2025)
por: Guo, Tianyu, et al.
Publicado: (2025)
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
por: Shen, Haiying, et al.
Publicado: (2025)
por: Shen, Haiying, et al.
Publicado: (2025)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
KV Cache Transform Coding for Compact Storage in LLM Inference
por: Staniszewski, Konrad, et al.
Publicado: (2025)
por: Staniszewski, Konrad, et al.
Publicado: (2025)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
por: Yang, Dongjie, et al.
Publicado: (2024)
por: Yang, Dongjie, et al.
Publicado: (2024)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
por: Gu, Yuzhe, et al.
Publicado: (2025)
por: Gu, Yuzhe, et al.
Publicado: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
por: Zhou, Yuhao, et al.
Publicado: (2025)
por: Zhou, Yuhao, et al.
Publicado: (2025)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
por: Wu, Haoyi, et al.
Publicado: (2024)
por: Wu, Haoyi, et al.
Publicado: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
por: Liao, Mengqi, et al.
Publicado: (2025)
por: Liao, Mengqi, et al.
Publicado: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
por: Patel, Ishan, et al.
Publicado: (2026)
por: Patel, Ishan, et al.
Publicado: (2026)
Ejemplares similares
-
Efficient Long-Context LLM Inference via KV Cache Clustering
por: Hu, Jie, et al.
Publicado: (2025) -
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
por: Tao, Wei, et al.
Publicado: (2026) -
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025) -
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024) -
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
por: Yang, Yifei, et al.
Publicado: (2024)