QAQ: Quality Adaptive Quantization for LLM KV Cache
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Shichen, Cheng, Wen, Qin, Jiayu, Wang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
von: Zuo, Youhui, et al.
Veröffentlicht: (2025)
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
von: Liu, Peiyu, et al.
Veröffentlicht: (2024)
von: Liu, Peiyu, et al.
Veröffentlicht: (2024)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
QAQ: Bidirectional Semantic Coherence for Selecting High-Quality Synthetic Code Instructions
von: Lei, Jiayin, et al.
Veröffentlicht: (2026)
von: Lei, Jiayin, et al.
Veröffentlicht: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
CommVQ: Commutative Vector Quantization for KV Cache Compression
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
von: Yu, Bohan, et al.
Veröffentlicht: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
Taming the Fragility of KV Cache Eviction in LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
von: Feng, Yuan, et al.
Veröffentlicht: (2025)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
von: Liu, Tengxuan, et al.
Veröffentlicht: (2025)
von: Liu, Tengxuan, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
von: Cai, Zhihang, et al.
Veröffentlicht: (2025)
von: Cai, Zhihang, et al.
Veröffentlicht: (2025)
dKV-Cache: The Cache for Diffusion Language Models
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
von: Ma, Xinyin, et al.
Veröffentlicht: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026)
von: Tao, Wei, et al.
Veröffentlicht: (2026)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
von: Chen, Jinhan, et al.
Veröffentlicht: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
von: An, Yuwei, et al.
Veröffentlicht: (2025)
von: An, Yuwei, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
von: Qin, Ziran, et al.
Veröffentlicht: (2025) -
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
von: Zuo, Youhui, et al.
Veröffentlicht: (2025) -
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025) -
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
von: Chen, Jian, et al.
Veröffentlicht: (2026) -
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025)