Practical Efficiency of Muon for Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915293049126912 |
|---|---|
| author | AI, Essential : Shah, Ishaan Polloreno, Anthony M. Stratos, Karl Monk, Philip Chaluvaraju, Adarsh Hojel, Andrew Ma, Andrew Thomas, Anil Tanwer, Ashish Shah, Darsh J Nguyen, Khoi Smith, Kurt Callahan, Michael Pust, Michael Parmar, Mohit Rushton, Peter Mazarakis, Platon Kapila, Ritvik Srivastava, Saurabh Singla, Somanshu Romanski, Tim Vanjani, Yash Vaswani, Ashish |
| author_facet | AI, Essential : Shah, Ishaan Polloreno, Anthony M. Stratos, Karl Monk, Philip Chaluvaraju, Adarsh Hojel, Andrew Ma, Andrew Thomas, Anil Tanwer, Ashish Shah, Darsh J Nguyen, Khoi Smith, Kurt Callahan, Michael Pust, Michael Parmar, Mohit Rushton, Peter Mazarakis, Platon Kapila, Ritvik Srivastava, Saurabh Singla, Somanshu Romanski, Tim Vanjani, Yash Vaswani, Ashish |
| contents | We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_02222 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Practical Efficiency of Muon for Pretraining AI, Essential : Shah, Ishaan Polloreno, Anthony M. Stratos, Karl Monk, Philip Chaluvaraju, Adarsh Hojel, Andrew Ma, Andrew Thomas, Anil Tanwer, Ashish Shah, Darsh J Nguyen, Khoi Smith, Kurt Callahan, Michael Pust, Michael Parmar, Mohit Rushton, Peter Mazarakis, Platon Kapila, Ritvik Srivastava, Saurabh Singla, Somanshu Romanski, Tim Vanjani, Yash Vaswani, Ashish Machine Learning We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture. |
| title | Practical Efficiency of Muon for Pretraining |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2505.02222 |