Muon Outperforms Adam in Tail-End Associative Memory Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Shuche, Zhang, Fengzhuo, Li, Jiaxiang, Du, Cunxiao, Du, Chao, Pang, Tianyu, Yang, Zhuoran, Hong, Mingyi, Tan, Vincent Y. F.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918154368712704
author Wang, Shuche
Zhang, Fengzhuo
Li, Jiaxiang
Du, Cunxiao
Du, Chao
Pang, Tianyu
Yang, Zhuoran
Hong, Mingyi
Tan, Vincent Y. F.
author_facet Wang, Shuche
Zhang, Fengzhuo
Li, Jiaxiang
Du, Cunxiao
Du, Chao
Pang, Tianyu
Yang, Zhuoran
Hong, Mingyi
Tan, Vincent Y. F.
contents The Muon optimizer is consistently faster than Adam in training Large Language Models (LLMs), yet the mechanism underlying its success remains unclear. This paper demystifies this mechanism through the lens of associative memory. By ablating the transformer components optimized by Muon, we reveal that the associative memory parameters of LLMs, namely the Value and Output (VO) attention weights and Feed-Forward Networks (FFNs), are the primary contributors to Muon's superiority. Motivated by this associative memory view, we then explain Muon's superiority on real-world corpora, which are intrinsically heavy-tailed: a few classes (tail classes) appear far less frequently than others. The superiority is explained through two key properties: (i) its update rule consistently yields a more isotropic singular spectrum than Adam; and as a result, (ii) on heavy-tailed data, it optimizes tail classes more effectively than Adam. Beyond empirical evidence, we theoretically confirm these findings by analyzing a one-layer associative memory model under class-imbalanced data. We prove that Muon consistently achieves balanced learning across classes regardless of feature embeddings, whereas Adam can induce large disparities in learning errors depending on embedding properties. In summary, our empirical observations and theoretical analyses reveal Muon's core advantage: its update rule aligns with the outer-product structure of linear associative memories, enabling more balanced and effective learning of tail classes in heavy-tailed distributions than Adam.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26030
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Muon Outperforms Adam in Tail-End Associative Memory Learning
Wang, Shuche
Zhang, Fengzhuo
Li, Jiaxiang
Du, Cunxiao
Du, Chao
Pang, Tianyu
Yang, Zhuoran
Hong, Mingyi
Tan, Vincent Y. F.
Machine Learning
Artificial Intelligence
Optimization and Control
The Muon optimizer is consistently faster than Adam in training Large Language Models (LLMs), yet the mechanism underlying its success remains unclear. This paper demystifies this mechanism through the lens of associative memory. By ablating the transformer components optimized by Muon, we reveal that the associative memory parameters of LLMs, namely the Value and Output (VO) attention weights and Feed-Forward Networks (FFNs), are the primary contributors to Muon's superiority. Motivated by this associative memory view, we then explain Muon's superiority on real-world corpora, which are intrinsically heavy-tailed: a few classes (tail classes) appear far less frequently than others. The superiority is explained through two key properties: (i) its update rule consistently yields a more isotropic singular spectrum than Adam; and as a result, (ii) on heavy-tailed data, it optimizes tail classes more effectively than Adam. Beyond empirical evidence, we theoretically confirm these findings by analyzing a one-layer associative memory model under class-imbalanced data. We prove that Muon consistently achieves balanced learning across classes regardless of feature embeddings, whereas Adam can induce large disparities in learning errors depending on embedding properties. In summary, our empirical observations and theoretical analyses reveal Muon's core advantage: its update rule aligns with the outer-product structure of linear associative memories, enabling more balanced and effective learning of tail classes in heavy-tailed distributions than Adam.
title Muon Outperforms Adam in Tail-End Associative Memory Learning
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2509.26030