RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Zuo, Fei, Zhou, Zikang, Cong, Hao, Xi, Xiaoyan, Leung, Ho Fai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rate-Distortion Theory for Mixed States
by: Khanian, Zahra Baghali, et al.
Published: (2022)
by: Khanian, Zahra Baghali, et al.
Published: (2022)
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
by: Park, Bumsu, et al.
Published: (2026)
by: Park, Bumsu, et al.
Published: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026)
by: Tao, Wei, et al.
Published: (2026)
Harmonizing Program Induction with Rate-Distortion Theory
by: Zhou, Hanqi, et al.
Published: (2024)
by: Zhou, Hanqi, et al.
Published: (2024)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
by: Lee, Namyoon, et al.
Published: (2026)
by: Lee, Namyoon, et al.
Published: (2026)
A Rate-Distortion Framework for Summarization
by: Arda, Enes, et al.
Published: (2025)
by: Arda, Enes, et al.
Published: (2025)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
by: Chauhan, Anay, et al.
Published: (2026)
by: Chauhan, Anay, et al.
Published: (2026)
AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
by: Leung, Ho Fai, et al.
Published: (2025)
by: Leung, Ho Fai, et al.
Published: (2025)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
by: Lu, Haiquan, et al.
Published: (2026)
by: Lu, Haiquan, et al.
Published: (2026)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
by: Chen, Han, et al.
Published: (2025)
by: Chen, Han, et al.
Published: (2025)
Semantic Rate-Distortion Theory with Applications
by: Zhao, Yi-Qun, et al.
Published: (2025)
by: Zhao, Yi-Qun, et al.
Published: (2025)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
by: Zhang, Junkai, et al.
Published: (2026)
by: Zhang, Junkai, et al.
Published: (2026)
Hallucination is a Consequence of Space-Optimality: A Rate-Distortion Theorem for Membership Testing
by: Guo, Anxin, et al.
Published: (2026)
by: Guo, Anxin, et al.
Published: (2026)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
by: Liu, Tengxuan, et al.
Published: (2025)
by: Liu, Tengxuan, et al.
Published: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
by: Rehg, Isaac
Published: (2024)
by: Rehg, Isaac
Published: (2024)
A Constrained BA Algorithm for Rate-Distortion and Distortion-Rate Functions
by: Chen, Lingyi, et al.
Published: (2023)
by: Chen, Lingyi, et al.
Published: (2023)
Statistical Inference and Quality Measures of KV Cache Quantisations Inspired by TurboQuant
by: D'Alberto, Paolo
Published: (2026)
by: D'Alberto, Paolo
Published: (2026)
Accurate KV Cache Quantization with Outlier Tokens Tracing
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
Stochastic Chase Decoding for BMS Channels via Rate Distortion Theory
by: Berman, Amit, et al.
Published: (2026)
by: Berman, Amit, et al.
Published: (2026)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
by: Nagle, Alliot, et al.
Published: (2024)
by: Nagle, Alliot, et al.
Published: (2024)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
by: Li, Xing, et al.
Published: (2025)
by: Li, Xing, et al.
Published: (2025)
Rate-Distortion-Classification Representation Theory for Bernoulli Sources
by: Nguyen, Nam, et al.
Published: (2026)
by: Nguyen, Nam, et al.
Published: (2026)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
RDD Function: A Tradeoff Between Rate and Distortion-in-Distortion
by: Chen, Lingyi, et al.
Published: (2025)
by: Chen, Lingyi, et al.
Published: (2025)
Optimal Neural Compressors for the Rate-Distortion-Perception Tradeoff
by: Lei, Eric, et al.
Published: (2025)
by: Lei, Eric, et al.
Published: (2025)
Gaussian Rate-Distortion-Perception Coding and Entropy-Constrained Scalar Quantization
by: Xie, Li, et al.
Published: (2024)
by: Xie, Li, et al.
Published: (2024)
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
by: Zhang, Jiebin, et al.
Published: (2024)
by: Zhang, Jiebin, et al.
Published: (2024)
QAQ: Quality Adaptive Quantization for LLM KV Cache
by: Dong, Shichen, et al.
Published: (2024)
by: Dong, Shichen, et al.
Published: (2024)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
by: Zandieh, Amir, et al.
Published: (2025)
by: Zandieh, Amir, et al.
Published: (2025)
Rate-Distortion-Perception Theory for the Quadratic Wasserstein Space
by: Qu, Xiqiang, et al.
Published: (2025)
by: Qu, Xiqiang, et al.
Published: (2025)
Characterizing the Optimal Memory-Rate Tradeoff in Secure Coded Caching for Small Buffer or Small Rate
by: Fang, Han, et al.
Published: (2025)
by: Fang, Han, et al.
Published: (2025)
Analyzing α-divergence in Gaussian Rate-Distortion-Perception Theory
by: Sourla, Martha V., et al.
Published: (2025)
by: Sourla, Martha V., et al.
Published: (2025)
Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees
by: Chen, Thomas Y.
Published: (2025)
by: Chen, Thomas Y.
Published: (2025)
FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels
by: Zuo, Fei, et al.
Published: (2026)
by: Zuo, Fei, et al.
Published: (2026)
IsoQuant: Hardware-Aligned SO(4) Isoclinic Rotations for LLM KV Cache Compression
by: Ji, Zhongping
Published: (2026)
by: Ji, Zhongping
Published: (2026)
Rate-Distortion Theory in Coding for Machines and its Application
by: Harell, Alon, et al.
Published: (2023)
by: Harell, Alon, et al.
Published: (2023)
Finite Block Length Rate-Distortion Theory for the Bernoulli Source with Hamming Distortion: A Tutorial
by: Krishnamachari, Bhaskar
Published: (2026)
by: Krishnamachari, Bhaskar
Published: (2026)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
by: Corallo, Giulio, et al.
Published: (2025)
by: Corallo, Giulio, et al.
Published: (2025)
Similar Items
-
Rate-Distortion Theory for Mixed States
by: Khanian, Zahra Baghali, et al.
Published: (2022) -
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
by: Park, Bumsu, et al.
Published: (2026) -
SQuat: Subspace-orthogonal KV Cache Quantization
by: Wang, Hao, et al.
Published: (2025) -
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026) -
Harmonizing Program Induction with Rate-Distortion Theory
by: Zhou, Hanqi, et al.
Published: (2024)