Saved in:
| Main Authors: | Kuzmin, Andrey, Van Baalen, Mart, Ren, Yuwei, Nagel, Markus, Peters, Jorn, Blankevoort, Tijmen |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2208.09225 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023)
by: Kuzmin, Andrey, et al.
Published: (2023)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024)
by: van Baalen, Mart, et al.
Published: (2024)
The LLM Surgeon
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
by: Skliar, Andrii, et al.
Published: (2024)
by: Skliar, Andrii, et al.
Published: (2024)
Rapid Switching and Multi-Adapter Fusion via Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
Sparse High Rank Adapters
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
by: Bhardwaj, Kartikeya, et al.
Published: (2024)
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
by: Cook, Jack, et al.
Published: (2025)
by: Cook, Jack, et al.
Published: (2025)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
by: Wang, Fengjuan, et al.
Published: (2025)
by: Wang, Fengjuan, et al.
Published: (2025)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
FPTQuant: Function-Preserving Transforms for LLM Quantization
by: van Breugel, Boris, et al.
Published: (2025)
by: van Breugel, Boris, et al.
Published: (2025)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)
by: Cao, Hengjie, et al.
Published: (2025)
Dissecting Quantization Error: A Concentration-Alignment Perspective
by: Federici, Marco, et al.
Published: (2026)
by: Federici, Marco, et al.
Published: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
by: Egiazarian, Vage, et al.
Published: (2025)
by: Egiazarian, Vage, et al.
Published: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
FireQ: Fast INT4-FP8 Kernel and RoPE-aware Quantization for LLM Inference Acceleration
by: Baek, Daehyeon, et al.
Published: (2025)
by: Baek, Daehyeon, et al.
Published: (2025)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
by: Liang, Guang, et al.
Published: (2025)
by: Liang, Guang, et al.
Published: (2025)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
by: Zhao, Maosen, et al.
Published: (2025)
by: Zhao, Maosen, et al.
Published: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
Reasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck
by: Massoli, Fabio Valerio, et al.
Published: (2026)
by: Massoli, Fabio Valerio, et al.
Published: (2026)
Towards Fully FP8 GEMM LLM Training at Scale
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Discovery of Decision Synchronization Patterns from Event Logs
by: Kuijpers, Tijmen, et al.
Published: (2026)
by: Kuijpers, Tijmen, et al.
Published: (2026)
When Flat Minima Fail: Characterizing INT4 Quantization Collapse After FP32 Convergence
by: Armstrong, Marcus
Published: (2026)
by: Armstrong, Marcus
Published: (2026)
$μ$nit Scaling: Simple and Scalable FP8 LLM Training
by: Narayan, Saaketh, et al.
Published: (2025)
by: Narayan, Saaketh, et al.
Published: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
by: Chen, Mengzhao, et al.
Published: (2025)
by: Chen, Mengzhao, et al.
Published: (2025)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
by: Xin, Meng, et al.
Published: (2026)
by: Xin, Meng, et al.
Published: (2026)
Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent Analysis
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Myosotis: structured computation for attention like layer
by: Egorov, Evgenii, et al.
Published: (2025)
by: Egorov, Evgenii, et al.
Published: (2025)
Similar Items
-
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023) -
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024) -
The LLM Surgeon
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023) -
Leech Lattice Vector Quantization for Efficient LLM Compression
by: van der Ouderaa, Tycho F. A., et al.
Published: (2026) -
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)