QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yanshu, Li, Wang, Yao, Zhaoqian, Yang, Tong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
por: Wang, Yanshu, et al.
Publicado: (2024)
por: Wang, Yanshu, et al.
Publicado: (2024)
An experimental study of KV cache reuse strategies in chunk-level caching systems
por: Cestola, Samuel, et al.
Publicado: (2026)
por: Cestola, Samuel, et al.
Publicado: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025)
por: Tian, Yuxuan, et al.
Publicado: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
por: Yu, Bohan, et al.
Publicado: (2025)
por: Yu, Bohan, et al.
Publicado: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
por: Yang, Qingyue, et al.
Publicado: (2025)
por: Yang, Qingyue, et al.
Publicado: (2025)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
por: Behnam, Payman, et al.
Publicado: (2025)
por: Behnam, Payman, et al.
Publicado: (2025)
Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization
por: Wang, YiFeng, et al.
Publicado: (2026)
por: Wang, YiFeng, et al.
Publicado: (2026)
Residual-Mass Accounting for Partial-KV Decoding
por: Hoshi, Yasuto, et al.
Publicado: (2026)
por: Hoshi, Yasuto, et al.
Publicado: (2026)
SALS: Sparse Attention in Latent Space for KV cache Compression
por: Mu, Junlin, et al.
Publicado: (2025)
por: Mu, Junlin, et al.
Publicado: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
por: Ye, Jiancai, et al.
Publicado: (2026)
por: Ye, Jiancai, et al.
Publicado: (2026)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
por: Chang, Chi-Chih, et al.
Publicado: (2025)
por: Chang, Chi-Chih, et al.
Publicado: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
por: Li, Xing, et al.
Publicado: (2025)
por: Li, Xing, et al.
Publicado: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
por: Jo, Dongwon, et al.
Publicado: (2025)
por: Jo, Dongwon, et al.
Publicado: (2025)
IsoQuant: Hardware-Aligned SO(4) Isoclinic Rotations for LLM KV Cache Compression
por: Ji, Zhongping
Publicado: (2026)
por: Ji, Zhongping
Publicado: (2026)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
por: Patel, Ishan, et al.
Publicado: (2026)
por: Patel, Ishan, et al.
Publicado: (2026)
LongFlow: Efficient KV Cache Compression for Reasoning Models
por: Su, Yi, et al.
Publicado: (2026)
por: Su, Yi, et al.
Publicado: (2026)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
por: Li, Weizhuo, et al.
Publicado: (2024)
por: Li, Weizhuo, et al.
Publicado: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
por: Wang, Hao, et al.
Publicado: (2025)
por: Wang, Hao, et al.
Publicado: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
por: Sun, Hanshi, et al.
Publicado: (2024)
por: Sun, Hanshi, et al.
Publicado: (2024)
Inference-Time Hyper-Scaling with KV Cache Compression
por: Łańcucki, Adrian, et al.
Publicado: (2025)
por: Łańcucki, Adrian, et al.
Publicado: (2025)
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
por: Li, Jiaqian, et al.
Publicado: (2026)
por: Li, Jiaqian, et al.
Publicado: (2026)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
por: Sharma, Akshat, et al.
Publicado: (2024)
por: Sharma, Akshat, et al.
Publicado: (2024)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
por: Lin, Bokai, et al.
Publicado: (2024)
por: Lin, Bokai, et al.
Publicado: (2024)
iFairy: the First 2-bit Complex LLM with All Parameters in $\{\pm1, \pm i\}$
por: Wang, Feiyu, et al.
Publicado: (2025)
por: Wang, Feiyu, et al.
Publicado: (2025)
KVSculpt: KV Cache Compression as Distillation
por: Jiang, Bo, et al.
Publicado: (2026)
por: Jiang, Bo, et al.
Publicado: (2026)
Residual vector quantization for KV cache compression in large language model
por: Kumar, Ankur
Publicado: (2024)
por: Kumar, Ankur
Publicado: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
por: Kang, Hao, et al.
Publicado: (2024)
por: Kang, Hao, et al.
Publicado: (2024)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
por: Saxena, Utkarsh, et al.
Publicado: (2024)
por: Saxena, Utkarsh, et al.
Publicado: (2024)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
por: Shao, Yuantian, et al.
Publicado: (2025)
por: Shao, Yuantian, et al.
Publicado: (2025)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
por: Vicentino, Caio
Publicado: (2026)
por: Vicentino, Caio
Publicado: (2026)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
por: Xiong, Boya, et al.
Publicado: (2025)
por: Xiong, Boya, et al.
Publicado: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
por: Yadav, Prateek, et al.
Publicado: (2023)
por: Yadav, Prateek, et al.
Publicado: (2023)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
por: Datta, Debajyoti, et al.
Publicado: (2026)
por: Datta, Debajyoti, et al.
Publicado: (2026)
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
por: Zeng, Hanqing, et al.
Publicado: (2025)
por: Zeng, Hanqing, et al.
Publicado: (2025)
Ejemplares similares
-
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
por: Wang, Yanshu, et al.
Publicado: (2024) -
An experimental study of KV cache reuse strategies in chunk-level caching systems
por: Cestola, Samuel, et al.
Publicado: (2026) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025) -
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026) -
EvolKV: Evolutionary KV Cache Compression for LLM Inference
por: Yu, Bohan, et al.
Publicado: (2025)