Pyramid Vector Quantization for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | van der Ouderaa, Tycho F. A., Croci, Maximilian L., Hilmkil, Agrin, Hensman, James |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
Noether's razor: Learning Conserved Quantities
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
Low-Rank Correction for Quantized LLMs
by: Scetbon, Meyer, et al.
Published: (2024)
by: Scetbon, Meyer, et al.
Published: (2024)
Variational Inference Failures Under Model Symmetries: Permutation Invariant Posteriors for Bayesian Neural Networks
by: Gelberg, Yoav, et al.
Published: (2024)
by: Gelberg, Yoav, et al.
Published: (2024)
AVID: Adapting Video Diffusion Models to World Models
by: Rigter, Marc, et al.
Published: (2024)
by: Rigter, Marc, et al.
Published: (2024)
A Fixed-Point Approach for Causal Generative Modeling
by: Scetbon, Meyer, et al.
Published: (2024)
by: Scetbon, Meyer, et al.
Published: (2024)
Amortized Inference of Causal Models via Conditional Fixed-Point Iterations
by: Mahajan, Divyat, et al.
Published: (2024)
by: Mahajan, Divyat, et al.
Published: (2024)
The LLM Surgeon
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
Towards Causal Foundation Model: on Duality between Causal Inference and Attention
by: Zhang, Jiaqi, et al.
Published: (2023)
by: Zhang, Jiaqi, et al.
Published: (2023)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
by: Xu, Yuzhuang, et al.
Published: (2024)
by: Xu, Yuzhuang, et al.
Published: (2024)
Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
by: Chen, Yung-Chin, et al.
Published: (2026)
by: Chen, Yung-Chin, et al.
Published: (2026)
DiSK: A Diffusion Model for Structured Knowledge
by: Kitouni, Ouail, et al.
Published: (2023)
by: Kitouni, Ouail, et al.
Published: (2023)
Online Vector Quantized Attention
by: Alonso, Nick, et al.
Published: (2026)
by: Alonso, Nick, et al.
Published: (2026)
Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
by: Xu, Jin, et al.
Published: (2026)
by: Xu, Jin, et al.
Published: (2026)
Masked Vector Quantization
by: Nguyen, David D., et al.
Published: (2023)
by: Nguyen, David D., et al.
Published: (2023)
Scalable Model-Based Clustering with Sequential Monte Carlo
by: Trojan, Connie, et al.
Published: (2026)
by: Trojan, Connie, et al.
Published: (2026)
Representation Collapsing Problems in Vector Quantization
by: Zhao, Wenhao, et al.
Published: (2024)
by: Zhao, Wenhao, et al.
Published: (2024)
Gaussian Mixture Vector Quantization with Aggregated Categorical Posterior
by: Yan, Mingyuan, et al.
Published: (2024)
by: Yan, Mingyuan, et al.
Published: (2024)
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
Robust Training of Vector Quantized Bottleneck Models
by: Łańcucki, Adrian, et al.
Published: (2020)
by: Łańcucki, Adrian, et al.
Published: (2020)
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
by: Zhai, Linwei, et al.
Published: (2026)
by: Zhai, Linwei, et al.
Published: (2026)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
A Vector-Quantized Foundation Model for Patient Behavior Monitoring
by: Oliver, Rodrigo, et al.
Published: (2025)
by: Oliver, Rodrigo, et al.
Published: (2025)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
by: Bang, Wonjun, et al.
Published: (2025)
by: Bang, Wonjun, et al.
Published: (2025)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
by: Zhao, Wenhao, et al.
Published: (2026)
by: Zhao, Wenhao, et al.
Published: (2026)
Generalized Radius and Integrated Codebook Transforms for Differentiable Vector Quantization
by: You, Haochen, et al.
Published: (2026)
by: You, Haochen, et al.
Published: (2026)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
by: Xu, Chen, et al.
Published: (2025)
by: Xu, Chen, et al.
Published: (2025)
Locally-Adaptive Quantization for Streaming Vector Search
by: Aguerrebere, Cecilia, et al.
Published: (2024)
by: Aguerrebere, Cecilia, et al.
Published: (2024)
4bit-Quantization in Vector-Embedding for RAG
by: Jeong, Taehee
Published: (2025)
by: Jeong, Taehee
Published: (2025)
Kernel $k$-Medoids as General Vector Quantization
by: Gerlach, Thore, et al.
Published: (2025)
by: Gerlach, Thore, et al.
Published: (2025)
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024)
by: Gupta, Tarun, et al.
Published: (2024)
Vector Quantization Prompting for Continual Learning
by: Jiao, Li, et al.
Published: (2024)
by: Jiao, Li, et al.
Published: (2024)
Restructuring Vector Quantization with the Rotation Trick
by: Fifty, Christopher, et al.
Published: (2024)
by: Fifty, Christopher, et al.
Published: (2024)
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
Improving Vector-Quantized Image Modeling with Latent Consistency-Matching Diffusion
by: Nguyen, Bac, et al.
Published: (2024)
by: Nguyen, Bac, et al.
Published: (2024)
DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
by: Vali, Mohammad Hassan, et al.
Published: (2025)
by: Vali, Mohammad Hassan, et al.
Published: (2025)
Similar Items
-
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026) -
Noether's razor: Learning Conserved Quantities
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024) -
Low-Rank Correction for Quantized LLMs
by: Scetbon, Meyer, et al.
Published: (2024) -
Variational Inference Failures Under Model Symmetries: Permutation Invariant Posteriors for Bayesian Neural Networks
by: Gelberg, Yoav, et al.
Published: (2024) -
AVID: Adapting Video Diffusion Models to World Models
by: Rigter, Marc, et al.
Published: (2024)