Muon Optimizer Accelerates Grokking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tveit, Amund, Remseth, Bjørn, Skogvold, Arve
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910916616912896
author Tveit, Amund
Remseth, Bjørn
Skogvold, Arve
author_facet Tveit, Amund
Remseth, Bjørn
Skogvold, Arve
contents This paper investigates the impact of different optimizers on the grokking phenomenon, where models exhibit delayed generalization. We conducted experiments across seven numerical tasks (primarily modular arithmetic) using a modern Transformer architecture. The experimental configuration systematically varied the optimizer (Muon vs. AdamW) and the softmax activation function (standard softmax, stablemax, and sparsemax) to assess their combined effect on learning dynamics. Our empirical evaluation reveals that the Muon optimizer, characterized by its use of spectral norm constraints and second-order information, significantly accelerates the onset of grokking compared to the widely used AdamW optimizer. Specifically, Muon reduced the mean grokking epoch from 153.09 to 102.89 across all configurations, a statistically significant difference (t = 5.0175, p = 6.33e-08). This suggests that the optimizer choice plays a crucial role in facilitating the transition from memorization to generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16041
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Muon Optimizer Accelerates Grokking
Tveit, Amund
Remseth, Bjørn
Skogvold, Arve
Machine Learning
Artificial Intelligence
I.2
This paper investigates the impact of different optimizers on the grokking phenomenon, where models exhibit delayed generalization. We conducted experiments across seven numerical tasks (primarily modular arithmetic) using a modern Transformer architecture. The experimental configuration systematically varied the optimizer (Muon vs. AdamW) and the softmax activation function (standard softmax, stablemax, and sparsemax) to assess their combined effect on learning dynamics. Our empirical evaluation reveals that the Muon optimizer, characterized by its use of spectral norm constraints and second-order information, significantly accelerates the onset of grokking compared to the widely used AdamW optimizer. Specifically, Muon reduced the mean grokking epoch from 153.09 to 102.89 across all configurations, a statistically significant difference (t = 5.0175, p = 6.33e-08). This suggests that the optimizer choice plays a crucial role in facilitating the transition from memorization to generalization.
title Muon Optimizer Accelerates Grokking
topic Machine Learning
Artificial Intelligence
I.2
url https://arxiv.org/abs/2504.16041