VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
Fuente:
arXiv
Salvato in:
| Autori principali: | Yao, Dingyu, Yang, Chenxu, Tong, Zhengyang, Lin, Zheng, Liu, Wei, Luan, Jian, Wang, Weiping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
di: Yao, Dingyu, et al.
Pubblicazione: (2025)
di: Yao, Dingyu, et al.
Pubblicazione: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
di: Yang, Haoqi, et al.
Pubblicazione: (2025)
di: Yang, Haoqi, et al.
Pubblicazione: (2025)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
di: Son, Donghyun, et al.
Pubblicazione: (2025)
di: Son, Donghyun, et al.
Pubblicazione: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
Accurate KV Cache Quantization with Outlier Tokens Tracing
di: Su, Yi, et al.
Pubblicazione: (2025)
di: Su, Yi, et al.
Pubblicazione: (2025)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
di: Yang, Dongjie, et al.
Pubblicazione: (2024)
di: Yang, Dongjie, et al.
Pubblicazione: (2024)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
di: Wu, Guanlong, et al.
Pubblicazione: (2026)
di: Wu, Guanlong, et al.
Pubblicazione: (2026)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
di: Tao, Qian, et al.
Pubblicazione: (2024)
di: Tao, Qian, et al.
Pubblicazione: (2024)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
di: Li, Zeyu, et al.
Pubblicazione: (2025)
di: Li, Zeyu, et al.
Pubblicazione: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
QAQ: Quality Adaptive Quantization for LLM KV Cache
di: Dong, Shichen, et al.
Pubblicazione: (2024)
di: Dong, Shichen, et al.
Pubblicazione: (2024)
Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
di: Zhang, Xi, et al.
Pubblicazione: (2025)
di: Zhang, Xi, et al.
Pubblicazione: (2025)
TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
di: Shen, Hanzhang, et al.
Pubblicazione: (2026)
di: Shen, Hanzhang, et al.
Pubblicazione: (2026)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
di: Boroujeni, Sayed Pedram Haeri, et al.
Pubblicazione: (2026)
di: Boroujeni, Sayed Pedram Haeri, et al.
Pubblicazione: (2026)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
di: Zhang, Junkai, et al.
Pubblicazione: (2026)
di: Zhang, Junkai, et al.
Pubblicazione: (2026)
CommVQ: Commutative Vector Quantization for KV Cache Compression
di: Li, Junyan, et al.
Pubblicazione: (2025)
di: Li, Junyan, et al.
Pubblicazione: (2025)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
di: Luo, Zhifan, et al.
Pubblicazione: (2025)
di: Luo, Zhifan, et al.
Pubblicazione: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
di: Li, Xing, et al.
Pubblicazione: (2025)
di: Li, Xing, et al.
Pubblicazione: (2025)
MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
di: Wang, Zongwu, et al.
Pubblicazione: (2025)
di: Wang, Zongwu, et al.
Pubblicazione: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
di: Xu, Yichun, et al.
Pubblicazione: (2026)
di: Xu, Yichun, et al.
Pubblicazione: (2026)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
di: Jia, Jinda, et al.
Pubblicazione: (2026)
di: Jia, Jinda, et al.
Pubblicazione: (2026)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
di: Zuo, Youhui, et al.
Pubblicazione: (2025)
di: Zuo, Youhui, et al.
Pubblicazione: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
di: Bang, Wonjun, et al.
Pubblicazione: (2025)
di: Bang, Wonjun, et al.
Pubblicazione: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025)
di: Yu, Bohan, et al.
Pubblicazione: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
di: Du, Dayou, et al.
Pubblicazione: (2025)
di: Du, Dayou, et al.
Pubblicazione: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
di: Zeng, Wenxuan, et al.
Pubblicazione: (2025)
di: Zeng, Wenxuan, et al.
Pubblicazione: (2025)
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
di: Su, Zunhai, et al.
Pubblicazione: (2025)
di: Su, Zunhai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
di: Yao, Dingyu, et al.
Pubblicazione: (2025) -
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
di: Su, Zunhai, et al.
Pubblicazione: (2025) -
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
di: Yang, Haoqi, et al.
Pubblicazione: (2025) -
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
di: Son, Donghyun, et al.
Pubblicazione: (2025) -
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)