Extreme Compression of Large Language Models via Additive Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Egiazarian, Vage, Panferov, Andrei, Kuznedelev, Denis, Frantar, Elias, Babenko, Artem, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026)
by: Egiazarian, Vage, et al.
Published: (2026)
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
by: Shutova, Alina, et al.
Published: (2025)
by: Shutova, Alina, et al.
Published: (2025)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)
by: Rodionov, Gleb, et al.
Published: (2025)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
by: Egiazarian, Vage, et al.
Published: (2025)
by: Egiazarian, Vage, et al.
Published: (2025)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)
by: Chen, Jiale, et al.
Published: (2025)
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
by: Frantar, Elias, et al.
Published: (2024)
by: Frantar, Elias, et al.
Published: (2024)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
by: Nrusimha, Aniruddha, et al.
Published: (2024)
by: Nrusimha, Aniruddha, et al.
Published: (2024)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
SPADE: Sparsity-Guided Debugging for Deep Neural Networks
by: Moakhar, Arshia Soltani, et al.
Published: (2023)
by: Moakhar, Arshia Soltani, et al.
Published: (2023)
AutoJudge: Judge Decoding Without Manual Annotation
by: Garipov, Roman, et al.
Published: (2025)
by: Garipov, Roman, et al.
Published: (2025)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and Beyond
by: Platonov, Oleg, et al.
Published: (2022)
by: Platonov, Oleg, et al.
Published: (2022)
On the Compressibility of Quantized Large Language Models
by: Mao, Yu, et al.
Published: (2024)
by: Mao, Yu, et al.
Published: (2024)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
A critical look at the evaluation of GNNs under heterophily: Are we really making progress?
by: Platonov, Oleg, et al.
Published: (2023)
by: Platonov, Oleg, et al.
Published: (2023)
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
by: Nicolicioiu, Armand, et al.
Published: (2024)
by: Nicolicioiu, Armand, et al.
Published: (2024)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
by: Young, Sean I.
Published: (2024)
by: Young, Sean I.
Published: (2024)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
by: Xu, Yuzhuang, et al.
Published: (2024)
by: Xu, Yuzhuang, et al.
Published: (2024)
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
by: Castro, Roberto L., et al.
Published: (2025)
by: Castro, Roberto L., et al.
Published: (2025)
LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit
by: Gong, Ruihao, et al.
Published: (2024)
by: Gong, Ruihao, et al.
Published: (2024)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
by: Nikdan, Mahdi, et al.
Published: (2024)
by: Nikdan, Mahdi, et al.
Published: (2024)
Quantization of Large Language Models with an Overdetermined Basis
by: Merkulov, Daniil, et al.
Published: (2024)
by: Merkulov, Daniil, et al.
Published: (2024)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)
by: Wu, Diyuan, et al.
Published: (2024)
KV Cache Offloading for Context-Intensive Tasks
by: Bocharnikov, Andrey, et al.
Published: (2026)
by: Bocharnikov, Andrey, et al.
Published: (2026)
Additive Large Language Models for Semi-Structured Text
by: K, Karthikeyan, et al.
Published: (2025)
by: K, Karthikeyan, et al.
Published: (2025)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
CompactifAI: Extreme Compression of Large Language Models using Quantum-Inspired Tensor Networks
by: Tomut, Andrei, et al.
Published: (2024)
by: Tomut, Andrei, et al.
Published: (2024)
Similar Items
-
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025) -
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026) -
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025) -
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
by: Shutova, Alina, et al.
Published: (2025) -
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)