Recipes for Pre-training LLMs with MXFP8

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mishra, Asit, Stosic, Dusan, Layton, Simon, Micikevicius, Paulius
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909741333086208
author Mishra, Asit
Stosic, Dusan
Layton, Simon
Micikevicius, Paulius
author_facet Mishra, Asit
Stosic, Dusan
Layton, Simon
Micikevicius, Paulius
contents Using fewer bits to represent model parameters and related tensors during pre-training has become a required technique for improving GPU efficiency without sacrificing accuracy. Microscaling (MX) formats introduced in NVIDIA Blackwell generation of GPUs represent a major advancement of this technique, making it practical to combine narrow floating-point data types with finer granularity per-block scaling factors. In turn, this enables both quantization of more tensors than previous approaches and more efficient execution of operations on those tensors. Effective use of MX-formats requires careful choices of various parameters. In this paper we review these choices and show how MXFP8-E4M3 datatype and a specific number conversion algorithm result in training sessions that match those carried out in BF16. We present results using models with up to 8B parameters, trained on high-quality datasets of up to 15T tokens.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recipes for Pre-training LLMs with MXFP8
Mishra, Asit
Stosic, Dusan
Layton, Simon
Micikevicius, Paulius
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Using fewer bits to represent model parameters and related tensors during pre-training has become a required technique for improving GPU efficiency without sacrificing accuracy. Microscaling (MX) formats introduced in NVIDIA Blackwell generation of GPUs represent a major advancement of this technique, making it practical to combine narrow floating-point data types with finer granularity per-block scaling factors. In turn, this enables both quantization of more tensors than previous approaches and more efficient execution of operations on those tensors. Effective use of MX-formats requires careful choices of various parameters. In this paper we review these choices and show how MXFP8-E4M3 datatype and a specific number conversion algorithm result in training sessions that match those carried out in BF16. We present results using models with up to 8B parameters, trained on high-quality datasets of up to 15T tokens.
title Recipes for Pre-training LLMs with MXFP8
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2506.08027