An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Crawshaw, Michael, Modi, Chirag, Liu, Mingrui, Gower, Robert M.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918158575599616
author Crawshaw, Michael
Modi, Chirag
Liu, Mingrui
Gower, Robert M.
author_facet Crawshaw, Michael
Modi, Chirag
Liu, Mingrui
Gower, Robert M.
contents To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization. We systematically explore different alternatives for aggregating norms across layers, both formalizing existing combinations of Adam and the recently proposed Muon as a type of non-Euclidean gradient descent, and deriving new variants of the Muon optimizer. Through a comprehensive experimental evaluation of the optimizers within our framework, we find that Muon is sensitive to the choice of learning rate, whereas a new variant we call MuonMax is significantly more robust. We then show how to combine any non-Euclidean gradient method with model based momentum (known as Momo). The new Momo variants of Muon are significantly more robust to hyperparameter tuning, and often achieve a better validation score. Thus for new tasks, where the optimal hyperparameters are not known, we advocate for using Momo in combination with MuonMax to save on costly hyperparameter tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
Crawshaw, Michael
Modi, Chirag
Liu, Mingrui
Gower, Robert M.
Machine Learning
To define a steepest descent method over a neural network, we need to choose a norm for each layer, a way to aggregate these norms across layers, and whether to use normalization. We systematically explore different alternatives for aggregating norms across layers, both formalizing existing combinations of Adam and the recently proposed Muon as a type of non-Euclidean gradient descent, and deriving new variants of the Muon optimizer. Through a comprehensive experimental evaluation of the optimizers within our framework, we find that Muon is sensitive to the choice of learning rate, whereas a new variant we call MuonMax is significantly more robust. We then show how to combine any non-Euclidean gradient method with model based momentum (known as Momo). The new Momo variants of Muon are significantly more robust to hyperparameter tuning, and often achieve a better validation score. Thus for new tasks, where the optimal hyperparameters are not known, we advocate for using Momo in combination with MuonMax to save on costly hyperparameter tuning.
title An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
topic Machine Learning
url https://arxiv.org/abs/2510.09827