Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Eschenhagen, Runa, Cai, Anna, Lee, Tsung-Hsien, Shi, Hao-Jun Michael
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915788187762688
author Eschenhagen, Runa
Cai, Anna
Lee, Tsung-Hsien
Shi, Hao-Jun Michael
author_facet Eschenhagen, Runa
Cai, Anna
Lee, Tsung-Hsien
Shi, Hao-Jun Michael
contents Optimizers leveraging the matrix structure in neural networks, such as Shampoo and Muon, are more data-efficient than element-wise algorithms like Adam and Signum. While in specific settings, Shampoo and Muon reduce to spectral descent analogous to how Adam and Signum reduce to sign descent, their general relationship and relative data efficiency under controlled settings remain unclear. Through extensive experiments on language models, we demonstrate that Shampoo achieves higher token efficiency than Muon, mirroring Adam's advantage over Signum. We show that Shampoo's update applied to weight matrices can be decomposed into an adapted Muon update. Consistent with this, Shampoo's benefits can be exclusively attributed to its application to weight matrices, challenging interpretations agnostic to parameter shapes. This admits a new perspective that also avoids shortcomings of related interpretations based on variance adaptation and whitening: rather than enforcing semi-orthogonality as in spectral descent, Shampoo's updates are time-averaged semi-orthogonal in expectation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09314
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
Eschenhagen, Runa
Cai, Anna
Lee, Tsung-Hsien
Shi, Hao-Jun Michael
Machine Learning
Artificial Intelligence
Optimizers leveraging the matrix structure in neural networks, such as Shampoo and Muon, are more data-efficient than element-wise algorithms like Adam and Signum. While in specific settings, Shampoo and Muon reduce to spectral descent analogous to how Adam and Signum reduce to sign descent, their general relationship and relative data efficiency under controlled settings remain unclear. Through extensive experiments on language models, we demonstrate that Shampoo achieves higher token efficiency than Muon, mirroring Adam's advantage over Signum. We show that Shampoo's update applied to weight matrices can be decomposed into an adapted Muon update. Consistent with this, Shampoo's benefits can be exclusively attributed to its application to weight matrices, challenging interpretations agnostic to parameter shapes. This admits a new perspective that also avoids shortcomings of related interpretations based on variance adaptation and whitening: rather than enforcing semi-orthogonality as in spectral descent, Shampoo's updates are time-averaged semi-orthogonal in expectation.
title Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.09314