VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yixuan, Shi, Qingyu, Zhou, Jiayu, Liu, Dianbo, He, Ziwei, Lin, Zhouhan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
por: Lu, Liming, et al.
Publicado: (2026)
por: Lu, Liming, et al.
Publicado: (2026)
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
por: Kai, Jushi, et al.
Publicado: (2025)
por: Kai, Jushi, et al.
Publicado: (2025)
CommVQ: Commutative Vector Quantization for KV Cache Compression
por: Li, Junyan, et al.
Publicado: (2025)
por: Li, Junyan, et al.
Publicado: (2025)
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025)
por: Zeng, Boyi, et al.
Publicado: (2025)
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
por: Zeng, Boyi, et al.
Publicado: (2025)
por: Zeng, Boyi, et al.
Publicado: (2025)
QAQ: Quality Adaptive Quantization for LLM KV Cache
por: Dong, Shichen, et al.
Publicado: (2024)
por: Dong, Shichen, et al.
Publicado: (2024)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
por: Zhao, Wenhao, et al.
Publicado: (2026)
por: Zhao, Wenhao, et al.
Publicado: (2026)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
por: Ji, Shiyu, et al.
Publicado: (2026)
por: Ji, Shiyu, et al.
Publicado: (2026)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
por: Shen, Junzhe, et al.
Publicado: (2026)
por: Shen, Junzhe, et al.
Publicado: (2026)
A$^2$ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization
por: He, Junhui, et al.
Publicado: (2025)
por: He, Junhui, et al.
Publicado: (2025)
TreeKV: Smooth Key-Value Cache Compression with Tree Structures
por: He, Ziwei, et al.
Publicado: (2025)
por: He, Ziwei, et al.
Publicado: (2025)
AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depth
por: Song, Shixiang, et al.
Publicado: (2026)
por: Song, Shixiang, et al.
Publicado: (2026)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025)
por: Zeng, Boyi, et al.
Publicado: (2025)
Balance of Number of Embedding and their Dimensions in Vector Quantization
por: Chen, Hang, et al.
Publicado: (2024)
por: Chen, Hang, et al.
Publicado: (2024)
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
por: Liu, Peiyu, et al.
Publicado: (2024)
por: Liu, Peiyu, et al.
Publicado: (2024)
xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
por: Chang, Chi-Chih, et al.
Publicado: (2025)
por: Chang, Chi-Chih, et al.
Publicado: (2025)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
por: Yang, Haoqi, et al.
Publicado: (2025)
por: Yang, Haoqi, et al.
Publicado: (2025)
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
por: He, Ziwei, et al.
Publicado: (2023)
por: He, Ziwei, et al.
Publicado: (2023)
Towards Controlled Table-to-Text Generation with Scientific Reasoning
por: Guo, Zhixin, et al.
Publicado: (2023)
por: Guo, Zhixin, et al.
Publicado: (2023)
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Masking
por: Li, He, et al.
Publicado: (2026)
por: Li, He, et al.
Publicado: (2026)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
por: Li, Zeyu, et al.
Publicado: (2025)
por: Li, Zeyu, et al.
Publicado: (2025)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
por: Wu, Shunlong, et al.
Publicado: (2026)
por: Wu, Shunlong, et al.
Publicado: (2026)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
por: Chen, Jian, et al.
Publicado: (2026)
por: Chen, Jian, et al.
Publicado: (2026)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
por: Yang, Dongjie, et al.
Publicado: (2024)
por: Yang, Dongjie, et al.
Publicado: (2024)
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
por: Wu, Jialong, et al.
Publicado: (2024)
por: Wu, Jialong, et al.
Publicado: (2024)
Accurate KV Cache Quantization with Outlier Tokens Tracing
por: Su, Yi, et al.
Publicado: (2025)
por: Su, Yi, et al.
Publicado: (2025)
OpenBA-V2: Reaching 77.3% High Compression Ratio with Fast Multi-Stage Pruning
por: Qiao, Dan, et al.
Publicado: (2024)
por: Qiao, Dan, et al.
Publicado: (2024)
Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
por: Liu, Xiaoran, et al.
Publicado: (2025)
por: Liu, Xiaoran, et al.
Publicado: (2025)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
por: Xu, Yuzhuang, et al.
Publicado: (2024)
por: Xu, Yuzhuang, et al.
Publicado: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
CaliDrop: KV Cache Compression with Calibration
por: Su, Yi, et al.
Publicado: (2025)
por: Su, Yi, et al.
Publicado: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
por: Shi, Zhiyuan, et al.
Publicado: (2026)
por: Shi, Zhiyuan, et al.
Publicado: (2026)
More Human, More Efficient: Aligning Annotations with Quantized SLMs
por: Wang, Jiayu, et al.
Publicado: (2026)
por: Wang, Jiayu, et al.
Publicado: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
por: Liu, Akide, et al.
Publicado: (2024)
por: Liu, Akide, et al.
Publicado: (2024)
Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMs
por: Cao, Zhiwei, et al.
Publicado: (2024)
por: Cao, Zhiwei, et al.
Publicado: (2024)
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
por: Lagzian, Arash, et al.
Publicado: (2025)
por: Lagzian, Arash, et al.
Publicado: (2025)
Ejemplares similares
-
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
por: Lu, Liming, et al.
Publicado: (2026) -
FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension
por: Kai, Jushi, et al.
Publicado: (2025) -
CommVQ: Commutative Vector Quantization for KV Cache Compression
por: Li, Junyan, et al.
Publicado: (2025) -
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
por: Zeng, Boyi, et al.
Publicado: (2025) -
AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
por: Zeng, Boyi, et al.
Publicado: (2025)