FP4 All the Way: Fully Quantized Training of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Chmiel, Brian, Fishman, Maxim, Banner, Ron, Soudry, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling FP8 training to trillion-token LLMs
por: Fishman, Maxim, et al.
Publicado: (2024)
por: Fishman, Maxim, et al.
Publicado: (2024)
Normalized Architectures are Natively 4-Bit
por: Fishman, Maxim, et al.
Publicado: (2026)
por: Fishman, Maxim, et al.
Publicado: (2026)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
por: Chmiel, Brian, et al.
Publicado: (2022)
por: Chmiel, Brian, et al.
Publicado: (2022)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
por: Shkolnik, Moran, et al.
Publicado: (2024)
por: Shkolnik, Moran, et al.
Publicado: (2024)
Workspace Optimization: How to Train Your Agent
por: Sarafian, Elad, et al.
Publicado: (2026)
por: Sarafian, Elad, et al.
Publicado: (2026)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
por: Chmiel, Brian, et al.
Publicado: (2021)
por: Chmiel, Brian, et al.
Publicado: (2021)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
por: Cao, Hengjie, et al.
Publicado: (2026)
por: Cao, Hengjie, et al.
Publicado: (2026)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
por: Zhao, Maosen, et al.
Publicado: (2025)
por: Zhao, Maosen, et al.
Publicado: (2025)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
por: Wang, Fengjuan, et al.
Publicado: (2025)
por: Wang, Fengjuan, et al.
Publicado: (2025)
Efficient Post-training Quantization with FP8 Formats
por: Shen, Haihao, et al.
Publicado: (2023)
por: Shen, Haihao, et al.
Publicado: (2023)
On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
por: Deutel, Mark, et al.
Publicado: (2024)
por: Deutel, Mark, et al.
Publicado: (2024)
Metis: Training LLMs with FP4 Quantization
por: Cao, Hengjie, et al.
Publicado: (2025)
por: Cao, Hengjie, et al.
Publicado: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
por: Zhou, Sifan, et al.
Publicado: (2025)
por: Zhou, Sifan, et al.
Publicado: (2025)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
por: Zhang, Wuyue, et al.
Publicado: (2026)
por: Zhang, Wuyue, et al.
Publicado: (2026)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
por: Lu, Jinming, et al.
Publicado: (2025)
por: Lu, Jinming, et al.
Publicado: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
por: Chen, Mengzhao, et al.
Publicado: (2025)
por: Chen, Mengzhao, et al.
Publicado: (2025)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
por: Zhang, Jinghe, et al.
Publicado: (2026)
por: Zhang, Jinghe, et al.
Publicado: (2026)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
por: Blumenfeld, Yaniv, et al.
Publicado: (2024)
por: Blumenfeld, Yaniv, et al.
Publicado: (2024)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
por: Xi, Haocheng, et al.
Publicado: (2024)
por: Xi, Haocheng, et al.
Publicado: (2024)
Achieving binary weight and activation for LLMs using Post-Training Quantization
por: Song, Siqing, et al.
Publicado: (2025)
por: Song, Siqing, et al.
Publicado: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
por: Tan, Qitao, et al.
Publicado: (2025)
por: Tan, Qitao, et al.
Publicado: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
por: Luo, Yingsong, et al.
Publicado: (2024)
por: Luo, Yingsong, et al.
Publicado: (2024)
Low-Rank Quantization-Aware Training for LLMs
por: Bondarenko, Yelysei, et al.
Publicado: (2024)
por: Bondarenko, Yelysei, et al.
Publicado: (2024)
Defeating the Training-Inference Mismatch via FP16
por: Qi, Penghui, et al.
Publicado: (2025)
por: Qi, Penghui, et al.
Publicado: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
por: Lee, Joonhyung, et al.
Publicado: (2024)
por: Lee, Joonhyung, et al.
Publicado: (2024)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
por: Zhao, Zhixiong, et al.
Publicado: (2026)
por: Zhao, Zhixiong, et al.
Publicado: (2026)
Pretraining large language models with MXFP4 on Native FP4 Hardware
por: Cim, Musa, et al.
Publicado: (2026)
por: Cim, Musa, et al.
Publicado: (2026)
Continuous Approximations for Improving Quantization Aware Training of LLMs
por: Li, He, et al.
Publicado: (2024)
por: Li, He, et al.
Publicado: (2024)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
por: Zhao, Jiaqi, et al.
Publicado: (2025)
por: Zhao, Jiaqi, et al.
Publicado: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
por: Yan, Xianglong, et al.
Publicado: (2026)
por: Yan, Xianglong, et al.
Publicado: (2026)
Attributions All the Way Down? The Metagame of Interpretability
por: Baniecki, Hubert, et al.
Publicado: (2026)
por: Baniecki, Hubert, et al.
Publicado: (2026)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
por: Zhang, Peiyuan, et al.
Publicado: (2026)
por: Zhang, Peiyuan, et al.
Publicado: (2026)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
por: Qiao, Dan, et al.
Publicado: (2024)
por: Qiao, Dan, et al.
Publicado: (2024)
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
por: Meng, Haoqian, et al.
Publicado: (2026)
por: Meng, Haoqian, et al.
Publicado: (2026)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
por: Hoffer, Elad, et al.
Publicado: (2026)
por: Hoffer, Elad, et al.
Publicado: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
por: Huang, Wei, et al.
Publicado: (2024)
por: Huang, Wei, et al.
Publicado: (2024)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
por: You, Jaeseong, et al.
Publicado: (2024)
por: You, Jaeseong, et al.
Publicado: (2024)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
por: Cho, Yoonjun, et al.
Publicado: (2026)
por: Cho, Yoonjun, et al.
Publicado: (2026)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
por: Wang, Guoan, et al.
Publicado: (2026)
por: Wang, Guoan, et al.
Publicado: (2026)
Ejemplares similares
-
Scaling FP8 training to trillion-token LLMs
por: Fishman, Maxim, et al.
Publicado: (2024) -
Normalized Architectures are Natively 4-Bit
por: Fishman, Maxim, et al.
Publicado: (2026) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
por: Chmiel, Brian, et al.
Publicado: (2022) -
EXAQ: Exponent Aware Quantization For LLMs Acceleration
por: Shkolnik, Moran, et al.
Publicado: (2024) -
Workspace Optimization: How to Train Your Agent
por: Sarafian, Elad, et al.
Publicado: (2026)