FP4 All the Way: Fully Quantized Training of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chmiel, Brian, Fishman, Maxim, Banner, Ron, Soudry, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling FP8 training to trillion-token LLMs
von: Fishman, Maxim, et al.
Veröffentlicht: (2024)
von: Fishman, Maxim, et al.
Veröffentlicht: (2024)
Normalized Architectures are Natively 4-Bit
von: Fishman, Maxim, et al.
Veröffentlicht: (2026)
von: Fishman, Maxim, et al.
Veröffentlicht: (2026)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
Workspace Optimization: How to Train Your Agent
von: Sarafian, Elad, et al.
Veröffentlicht: (2026)
von: Sarafian, Elad, et al.
Veröffentlicht: (2026)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
von: Chmiel, Brian, et al.
Veröffentlicht: (2021)
von: Chmiel, Brian, et al.
Veröffentlicht: (2021)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
von: Zhao, Maosen, et al.
Veröffentlicht: (2025)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
von: Wang, Fengjuan, et al.
Veröffentlicht: (2025)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
von: Deutel, Mark, et al.
Veröffentlicht: (2024)
von: Deutel, Mark, et al.
Veröffentlicht: (2024)
Metis: Training LLMs with FP4 Quantization
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)
von: Cao, Hengjie, et al.
Veröffentlicht: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
von: Zhou, Sifan, et al.
Veröffentlicht: (2025)
von: Zhou, Sifan, et al.
Veröffentlicht: (2025)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wuyue, et al.
Veröffentlicht: (2026)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghe, et al.
Veröffentlicht: (2026)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
von: Xi, Haocheng, et al.
Veröffentlicht: (2024)
Achieving binary weight and activation for LLMs using Post-Training Quantization
von: Song, Siqing, et al.
Veröffentlicht: (2025)
von: Song, Siqing, et al.
Veröffentlicht: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
von: Luo, Yingsong, et al.
Veröffentlicht: (2024)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2026)
Pretraining large language models with MXFP4 on Native FP4 Hardware
von: Cim, Musa, et al.
Veröffentlicht: (2026)
von: Cim, Musa, et al.
Veröffentlicht: (2026)
Continuous Approximations for Improving Quantization Aware Training of LLMs
von: Li, He, et al.
Veröffentlicht: (2024)
von: Li, He, et al.
Veröffentlicht: (2024)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
Attributions All the Way Down? The Metagame of Interpretability
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
von: Qiao, Dan, et al.
Veröffentlicht: (2024)
von: Qiao, Dan, et al.
Veröffentlicht: (2024)
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
von: Meng, Haoqian, et al.
Veröffentlicht: (2026)
von: Meng, Haoqian, et al.
Veröffentlicht: (2026)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
von: Hoffer, Elad, et al.
Veröffentlicht: (2026)
von: Hoffer, Elad, et al.
Veröffentlicht: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
von: You, Jaeseong, et al.
Veröffentlicht: (2024)
von: You, Jaeseong, et al.
Veröffentlicht: (2024)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
von: Cho, Yoonjun, et al.
Veröffentlicht: (2026)
von: Cho, Yoonjun, et al.
Veröffentlicht: (2026)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
von: Wang, Guoan, et al.
Veröffentlicht: (2026)
von: Wang, Guoan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling FP8 training to trillion-token LLMs
von: Fishman, Maxim, et al.
Veröffentlicht: (2024) -
Normalized Architectures are Natively 4-Bit
von: Fishman, Maxim, et al.
Veröffentlicht: (2026) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022) -
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024) -
Workspace Optimization: How to Train Your Agent
von: Sarafian, Elad, et al.
Veröffentlicht: (2026)