Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Fujii, Kazuki, Nakamura, Taishi, Yokota, Rio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026)
by: Nohara, Daisuke, et al.
Published: (2026)
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
by: Li, Yitong, et al.
Published: (2026)
by: Li, Yitong, et al.
Published: (2026)
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)
by: Cao, Hengjie, et al.
Published: (2025)
Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
by: Fujii, Kazuki, et al.
Published: (2025)
by: Fujii, Kazuki, et al.
Published: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
Towards Fully FP8 GEMM LLM Training at Scale
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Defeating the Training-Inference Mismatch via FP16
by: Qi, Penghui, et al.
Published: (2025)
by: Qi, Penghui, et al.
Published: (2025)
$μ$nit Scaling: Simple and Scalable FP8 LLM Training
by: Narayan, Saaketh, et al.
Published: (2025)
by: Narayan, Saaketh, et al.
Published: (2025)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
by: Xi, Haocheng, et al.
Published: (2024)
by: Xi, Haocheng, et al.
Published: (2024)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Stability and Accuracy Trade-offs in Statistical Estimation
by: Chakraborty, Abhinav, et al.
Published: (2026)
by: Chakraborty, Abhinav, et al.
Published: (2026)
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning
by: Raghavendra, Mohit, et al.
Published: (2025)
by: Raghavendra, Mohit, et al.
Published: (2025)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
by: Wang, Fengjuan, et al.
Published: (2025)
by: Wang, Fengjuan, et al.
Published: (2025)
Muon: Training and Trade-offs with Latent Attention and MoE
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
Understanding Generalization of Federated Learning: the Trade-off between Model Stability and Optimization
by: Zeng, Dun, et al.
Published: (2024)
by: Zeng, Dun, et al.
Published: (2024)
IBCL: Zero-shot Model Generation under Stability-Plasticity Trade-offs
by: Lu, Pengyuan, et al.
Published: (2023)
by: Lu, Pengyuan, et al.
Published: (2023)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
by: Ishikawa, Satoki, et al.
Published: (2024)
by: Ishikawa, Satoki, et al.
Published: (2024)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
Variational Learning Finds Flatter Solutions at the Edge of Stability
by: Ghosh, Avrajit, et al.
Published: (2025)
by: Ghosh, Avrajit, et al.
Published: (2025)
The Impact of Anisotropic Covariance Structure on the Training Dynamics and Generalization Error of Linear Networks
by: Watanabe, Taishi, et al.
Published: (2026)
by: Watanabe, Taishi, et al.
Published: (2026)
On the Relationship Between Double Descent of CNNs and Shape/Texture Bias Under Learning Process
by: Iwase, Shun, et al.
Published: (2025)
by: Iwase, Shun, et al.
Published: (2025)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
by: Castro, Roberto L., et al.
Published: (2025)
by: Castro, Roberto L., et al.
Published: (2025)
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
by: Wei, Wang, et al.
Published: (2025)
by: Wei, Wang, et al.
Published: (2025)
Mathematical models for off-ball scoring prediction in basketball
by: Kono, Rikako, et al.
Published: (2024)
by: Kono, Rikako, et al.
Published: (2024)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
by: Liang, Guang, et al.
Published: (2025)
by: Liang, Guang, et al.
Published: (2025)
Do Heavy Tails Help Diffusion? On the Subtle Trade-off Between Initialization and Training
by: Cherkaoui, Hamza, et al.
Published: (2026)
by: Cherkaoui, Hamza, et al.
Published: (2026)
Misclassification Rate and Privacy-Utility Trade-offs in Graph Convolutional Networks via Subsampling Stability
by: Zhang, Yexin, et al.
Published: (2026)
by: Zhang, Yexin, et al.
Published: (2026)
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
by: Cao, Hengjie, et al.
Published: (2026)
by: Cao, Hengjie, et al.
Published: (2026)
Similar Items
-
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025) -
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026) -
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
by: Fujii, Kazuki, et al.
Published: (2024) -
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
by: Li, Yitong, et al.
Published: (2026) -
Metis: Training LLMs with FP4 Quantization
by: Cao, Hengjie, et al.
Published: (2025)