The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gronich, Eitan, Vardi, Gal
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916042017603584
author Gronich, Eitan
Vardi, Gal
author_facet Gronich, Eitan
Vardi, Gal
contents We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that \textit{momentum steepest descent} algorithms like Muon (spectral norm), MomentumGD ($\ell_2$ norm), and Signum ($\ell_\infty$ norm) are \textit{approximate} steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the $\ell_\infty$ margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16340
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
Gronich, Eitan
Vardi, Gal
Machine Learning
We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that \textit{momentum steepest descent} algorithms like Muon (spectral norm), MomentumGD ($\ell_2$ norm), and Signum ($\ell_\infty$ norm) are \textit{approximate} steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the $\ell_\infty$ margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.
title The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
topic Machine Learning
url https://arxiv.org/abs/2602.16340