HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Williams, Jorge L. Ruiz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
von: Zhang, Junkai, et al.
Veröffentlicht: (2026)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
von: Tao, Qian, et al.
Veröffentlicht: (2024)
von: Tao, Qian, et al.
Veröffentlicht: (2024)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
TurboAngle: Near-Lossless KV Cache Compression via Uniform Angle Quantization
von: Patel, Dipkumar
Veröffentlicht: (2026)
von: Patel, Dipkumar
Veröffentlicht: (2026)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
KV Cache Quantization for Self-Forcing Video Generation: A 33-Method Empirical Study
von: Ranganath, Suraj, et al.
Veröffentlicht: (2026)
von: Ranganath, Suraj, et al.
Veröffentlicht: (2026)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
von: Cai, Zhihang, et al.
Veröffentlicht: (2025)
von: Cai, Zhihang, et al.
Veröffentlicht: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
von: Son, Donghyun, et al.
Veröffentlicht: (2025)
von: Son, Donghyun, et al.
Veröffentlicht: (2025)
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
von: Xia, Haojun, et al.
Veröffentlicht: (2025)
von: Xia, Haojun, et al.
Veröffentlicht: (2025)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
von: Yang, Bin, et al.
Veröffentlicht: (2025)
von: Yang, Bin, et al.
Veröffentlicht: (2025)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
von: Chen, Kaiwen, et al.
Veröffentlicht: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
PatternKV: Flattening KV Representation Expands Quantization Headroom
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion
von: Tuncer, Tuna, et al.
Veröffentlicht: (2026)
von: Tuncer, Tuna, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
von: Zhang, Junkai, et al.
Veröffentlicht: (2026) -
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
von: Tao, Qian, et al.
Veröffentlicht: (2024) -
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026) -
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025) -
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
von: He, Yefei, et al.
Veröffentlicht: (2024)