Recipes for Pre-training LLMs with MXFP8
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909741333086208 |
|---|---|
| author | Mishra, Asit Stosic, Dusan Layton, Simon Micikevicius, Paulius |
| author_facet | Mishra, Asit Stosic, Dusan Layton, Simon Micikevicius, Paulius |
| contents | Using fewer bits to represent model parameters and related tensors during pre-training has become a required technique for improving GPU efficiency without sacrificing accuracy. Microscaling (MX) formats introduced in NVIDIA Blackwell generation of GPUs represent a major advancement of this technique, making it practical to combine narrow floating-point data types with finer granularity per-block scaling factors. In turn, this enables both quantization of more tensors than previous approaches and more efficient execution of operations on those tensors.
Effective use of MX-formats requires careful choices of various parameters. In this paper we review these choices and show how MXFP8-E4M3 datatype and a specific number conversion algorithm result in training sessions that match those carried out in BF16. We present results using models with up to 8B parameters, trained on high-quality datasets of up to 15T tokens. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_08027 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Recipes for Pre-training LLMs with MXFP8 Mishra, Asit Stosic, Dusan Layton, Simon Micikevicius, Paulius Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing Using fewer bits to represent model parameters and related tensors during pre-training has become a required technique for improving GPU efficiency without sacrificing accuracy. Microscaling (MX) formats introduced in NVIDIA Blackwell generation of GPUs represent a major advancement of this technique, making it practical to combine narrow floating-point data types with finer granularity per-block scaling factors. In turn, this enables both quantization of more tensors than previous approaches and more efficient execution of operations on those tensors. Effective use of MX-formats requires careful choices of various parameters. In this paper we review these choices and show how MXFP8-E4M3 datatype and a specific number conversion algorithm result in training sessions that match those carried out in BF16. We present results using models with up to 8B parameters, trained on high-quality datasets of up to 15T tokens. |
| title | Recipes for Pre-training LLMs with MXFP8 |
| topic | Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2506.08027 |