Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Egiazarian, Vage, Castro, Roberto L., Kuznedelev, Denis, Panferov, Andrei, Kurtic, Eldar, Pandit, Shubhra, Marques, Alexandre, Kurtz, Mark, Ashkboos, Saleh, Hoefler, Torsten, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026)
by: Egiazarian, Vage, et al.
Published: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)
by: Chen, Jiale, et al.
Published: (2025)
Beyond Outliers: A Study of Optimizers Under Quantization
by: Vlassis, Georgios, et al.
Published: (2025)
by: Vlassis, Georgios, et al.
Published: (2025)
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
by: Castro, Roberto L., et al.
Published: (2025)
by: Castro, Roberto L., et al.
Published: (2025)
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
by: Shutova, Alina, et al.
Published: (2025)
by: Shutova, Alina, et al.
Published: (2025)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
EfQAT: An Efficient Framework for Quantization-Aware Training
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
by: Agarwalla, Abhinav, et al.
Published: (2024)
by: Agarwalla, Abhinav, et al.
Published: (2024)
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
by: Ashkboos, Saleh, et al.
Published: (2025)
by: Ashkboos, Saleh, et al.
Published: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
by: Nicolicioiu, Armand, et al.
Published: (2024)
by: Nicolicioiu, Armand, et al.
Published: (2024)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)
by: Rodionov, Gleb, et al.
Published: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
by: Lin, Haokun, et al.
Published: (2026)
by: Lin, Haokun, et al.
Published: (2026)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
by: Chen, Jiale, et al.
Published: (2025)
by: Chen, Jiale, et al.
Published: (2025)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
Correlated Quantization for Faster Nonconvex Distributed Optimization
by: Panferov, Andrei, et al.
Published: (2024)
by: Panferov, Andrei, et al.
Published: (2024)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)
by: Wu, Diyuan, et al.
Published: (2024)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
by: Frantar, Elias, et al.
Published: (2024)
by: Frantar, Elias, et al.
Published: (2024)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
by: Malinovskii, Vladimir, et al.
Published: (2024)
by: Malinovskii, Vladimir, et al.
Published: (2024)
Neural Optimal Transport with General Cost Functionals
by: Asadulaev, Arip, et al.
Published: (2022)
by: Asadulaev, Arip, et al.
Published: (2022)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
by: Li, Shigang, et al.
Published: (2019)
by: Li, Shigang, et al.
Published: (2019)
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
by: Gordon, Ofir, et al.
Published: (2026)
by: Gordon, Ofir, et al.
Published: (2026)
Post Training Quantization of Large Language Models with Microscaling Formats
by: Sharify, Sayeh, et al.
Published: (2024)
by: Sharify, Sayeh, et al.
Published: (2024)
Similar Items
-
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024) -
Extreme Compression of Large Language Models via Additive Quantization
by: Egiazarian, Vage, et al.
Published: (2024) -
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
by: Egiazarian, Vage, et al.
Published: (2026) -
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024) -
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
by: Chen, Jiale, et al.
Published: (2025)