MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiacheng, Tan, Jianchao, Xu, Hongtao, Zhang, Jiaqi, Lu, Yifan, Sun, Yerui, Xie, Yuchen, Cai, Xunliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913163992104960
author Li, Jiacheng
Tan, Jianchao
Xu, Hongtao
Zhang, Jiaqi
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
author_facet Li, Jiacheng
Tan, Jianchao
Xu, Hongtao
Zhang, Jiaqi
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
contents The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term enables escape from sharp minima while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26842
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
Li, Jiacheng
Tan, Jianchao
Xu, Hongtao
Zhang, Jiaqi
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
Machine Learning
Computation and Language
The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an optimizer that bridges Muon's orthogonalization framework with curvature-aware acceleration. MONA adds an acceleration term directly into Muon's gradient processing pipeline. This term is calculated from the exponential moving average of gradient differences. We provide a detailed convergence analysis for MONA, showing that the acceleration term enables escape from sharp minima while preserving Muon's spectral-norm regularization. Empirically, MONA achieves better convergence and downstream task performance compared to both Muon and AdamW across three scales of Mixture-of-Experts pretraining, spanning from 1B to 68B parameters, with the largest model trained on 1 trillion tokens. Furthermore, we conduct supervised fine-tuning on the MOE-68B-A3B model and evaluate it on general capability, mathematical reasoning, and code generation benchmarks, where MONA achieves SOTA performance.
title MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.26842