Muon is Not That Special: Random or Inverted Spectra Work Just as Well

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shumaylov, Zakhar, Da Costa, Nathaël, Zaika, Peter, Mucsányi, Bálint, Massucco, Alex, Gelberg, Yoav, Schönlieb, Carola-Bibiane, Gal, Yarin, Hennig, Philipp
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910209874591744
author Shumaylov, Zakhar
Da Costa, Nathaël
Zaika, Peter
Mucsányi, Bálint
Massucco, Alex
Gelberg, Yoav
Schönlieb, Carola-Bibiane
Gal, Yarin
Hennig, Philipp
author_facet Shumaylov, Zakhar
Da Costa, Nathaël
Zaika, Peter
Mucsányi, Bálint
Massucco, Alex
Gelberg, Yoav
Schönlieb, Carola-Bibiane
Gal, Yarin
Hennig, Philipp
contents The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear minimization oracle (LMO) theory. In this paper, we challenge this geometric narrative through three contributions, demonstrating that precise geometric structure is not the key factor affecting optimization performance. First, we introduce Freon, a family of optimizers based on Schatten (quasi-)norms, powered by a novel, provably optimal QDWH-based iterative approximation. Freon naturally interpolates between SGD and Muon, while smoothly extrapolating into the quasi-norm regime. Empirically, the best-performing Schatten parameters for GPT-2 lie strictly within the quasi-norm regime, and thus cannot be represented by any unitarily invariant LMO. Second, noting that Freon performs well across a wide range of exponents, we introduce Kaon, an absurd optimizer that replaces singular values with random noise. Despite lacking any coherent geometric structure, Kaon matches Muon's performance and retains classical convergence guarantees, proving that strict adherence to a precise geometry is practically irrelevant. Third, having shown that geometry is not the primary driver of performance, we demonstrate it is instead controlled by two local quantities: alignment and descent potential. Ultimately, each optimizer must tune its step size around these two quantities. While their dynamics are difficult to predict a-priori, evaluating them within a stochastic random feature model yields a precise insight: Muon succeeds not by tracking an ideal global geometry, but by guaranteeing step-size optimality.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11181
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Muon is Not That Special: Random or Inverted Spectra Work Just as Well
Shumaylov, Zakhar
Da Costa, Nathaël
Zaika, Peter
Mucsányi, Bálint
Massucco, Alex
Gelberg, Yoav
Schönlieb, Carola-Bibiane
Gal, Yarin
Hennig, Philipp
Machine Learning
Artificial Intelligence
Numerical Analysis
Optimization and Control
The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear minimization oracle (LMO) theory. In this paper, we challenge this geometric narrative through three contributions, demonstrating that precise geometric structure is not the key factor affecting optimization performance. First, we introduce Freon, a family of optimizers based on Schatten (quasi-)norms, powered by a novel, provably optimal QDWH-based iterative approximation. Freon naturally interpolates between SGD and Muon, while smoothly extrapolating into the quasi-norm regime. Empirically, the best-performing Schatten parameters for GPT-2 lie strictly within the quasi-norm regime, and thus cannot be represented by any unitarily invariant LMO. Second, noting that Freon performs well across a wide range of exponents, we introduce Kaon, an absurd optimizer that replaces singular values with random noise. Despite lacking any coherent geometric structure, Kaon matches Muon's performance and retains classical convergence guarantees, proving that strict adherence to a precise geometry is practically irrelevant. Third, having shown that geometry is not the primary driver of performance, we demonstrate it is instead controlled by two local quantities: alignment and descent potential. Ultimately, each optimizer must tune its step size around these two quantities. While their dynamics are difficult to predict a-priori, evaluating them within a stochastic random feature model yields a precise insight: Muon succeeds not by tracking an ideal global geometry, but by guaranteeing step-size optimality.
title Muon is Not That Special: Random or Inverted Spectra Work Just as Well
topic Machine Learning
Artificial Intelligence
Numerical Analysis
Optimization and Control
url https://arxiv.org/abs/2605.11181