AdaMuon: Adaptive Muon Optimizer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Si, Chongjie, Zhang, Debing, Shen, Wei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911336636612608
author Si, Chongjie
Zhang, Debing
Shen, Wei
author_facet Si, Chongjie
Zhang, Debing
Shen, Wei
contents We propose AdaMuon, a novel optimizer that combines element-wise adaptivity with orthogonal updates for large-scale neural network training. AdaMuon incorporates two tightly coupled mechanisms: (1) an element-wise second momentum estimator applied to orthogonalized update directions, and (2) a sign-stabilized orthogonal update, where the momentum is first sign-transformed before orthogonalization. These two components jointly enable variance-adaptive scaling while maintaining stable update geometry. In addition, AdaMuon employs an RMS-aligned rescaling strategy to match the root-mean-square update magnitude to Adam, allowing direct reuse of existing learning rate schedules without extra tuning. Experiments demonstrate that AdaMuon not only maintains stability but can surpass Adam by more than 40\% training efficiency in large-scale scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaMuon: Adaptive Muon Optimizer
Si, Chongjie
Zhang, Debing
Shen, Wei
Machine Learning
We propose AdaMuon, a novel optimizer that combines element-wise adaptivity with orthogonal updates for large-scale neural network training. AdaMuon incorporates two tightly coupled mechanisms: (1) an element-wise second momentum estimator applied to orthogonalized update directions, and (2) a sign-stabilized orthogonal update, where the momentum is first sign-transformed before orthogonalization. These two components jointly enable variance-adaptive scaling while maintaining stable update geometry. In addition, AdaMuon employs an RMS-aligned rescaling strategy to match the root-mean-square update magnitude to Adam, allowing direct reuse of existing learning rate schedules without extra tuning. Experiments demonstrate that AdaMuon not only maintains stability but can surpass Adam by more than 40\% training efficiency in large-scale scenarios.
title AdaMuon: Adaptive Muon Optimizer
topic Machine Learning
url https://arxiv.org/abs/2507.11005