Practical Efficiency of Muon for Pretraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: AI, Essential, :, Shah, Ishaan, Polloreno, Anthony M., Stratos, Karl, Monk, Philip, Chaluvaraju, Adarsh, Hojel, Andrew, Ma, Andrew, Thomas, Anil, Tanwer, Ashish, Shah, Darsh J, Nguyen, Khoi, Smith, Kurt, Callahan, Michael, Pust, Michael, Parmar, Mohit, Rushton, Peter, Mazarakis, Platon, Kapila, Ritvik, Srivastava, Saurabh, Singla, Somanshu, Romanski, Tim, Vanjani, Yash, Vaswani, Ashish
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915293049126912
author AI, Essential
:
Shah, Ishaan
Polloreno, Anthony M.
Stratos, Karl
Monk, Philip
Chaluvaraju, Adarsh
Hojel, Andrew
Ma, Andrew
Thomas, Anil
Tanwer, Ashish
Shah, Darsh J
Nguyen, Khoi
Smith, Kurt
Callahan, Michael
Pust, Michael
Parmar, Mohit
Rushton, Peter
Mazarakis, Platon
Kapila, Ritvik
Srivastava, Saurabh
Singla, Somanshu
Romanski, Tim
Vanjani, Yash
Vaswani, Ashish
author_facet AI, Essential
:
Shah, Ishaan
Polloreno, Anthony M.
Stratos, Karl
Monk, Philip
Chaluvaraju, Adarsh
Hojel, Andrew
Ma, Andrew
Thomas, Anil
Tanwer, Ashish
Shah, Darsh J
Nguyen, Khoi
Smith, Kurt
Callahan, Michael
Pust, Michael
Parmar, Mohit
Rushton, Peter
Mazarakis, Platon
Kapila, Ritvik
Srivastava, Saurabh
Singla, Somanshu
Romanski, Tim
Vanjani, Yash
Vaswani, Ashish
contents We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Practical Efficiency of Muon for Pretraining
AI, Essential
:
Shah, Ishaan
Polloreno, Anthony M.
Stratos, Karl
Monk, Philip
Chaluvaraju, Adarsh
Hojel, Andrew
Ma, Andrew
Thomas, Anil
Tanwer, Ashish
Shah, Darsh J
Nguyen, Khoi
Smith, Kurt
Callahan, Michael
Pust, Michael
Parmar, Mohit
Rushton, Peter
Mazarakis, Platon
Kapila, Ritvik
Srivastava, Saurabh
Singla, Somanshu
Romanski, Tim
Vanjani, Yash
Vaswani, Ashish
Machine Learning
We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture.
title Practical Efficiency of Muon for Pretraining
topic Machine Learning
url https://arxiv.org/abs/2505.02222