Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Hengshuai, Chen, Xing, Murtadha, Ahmed, Wang, Guan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GAIN: Multiplicative Modulation for Domain Adaptation
von: Yao, Hengshuai, et al.
Veröffentlicht: (2026)
von: Yao, Hengshuai, et al.
Veröffentlicht: (2026)
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Beyond KV Caching: Shared Attention for Efficient LLMs
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
von: Liao, Bingli, et al.
Veröffentlicht: (2024)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
von: Yang, Dongquan, et al.
Veröffentlicht: (2025)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
von: Dong, Yanhao, et al.
Veröffentlicht: (2025)
von: Dong, Yanhao, et al.
Veröffentlicht: (2025)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
von: Yang, Bin, et al.
Veröffentlicht: (2025)
von: Yang, Bin, et al.
Veröffentlicht: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Why Attend to Everything? Focus is the Key
von: Yao, Hengshuai, et al.
Veröffentlicht: (2026)
von: Yao, Hengshuai, et al.
Veröffentlicht: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
von: Son, Donghyun, et al.
Veröffentlicht: (2025)
von: Son, Donghyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GAIN: Multiplicative Modulation for Domain Adaptation
von: Yao, Hengshuai, et al.
Veröffentlicht: (2026) -
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025) -
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024) -
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024) -
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)