Muon Outperforms Adam in Tail-End Associative Memory Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shuche, Zhang, Fengzhuo, Li, Jiaxiang, Du, Cunxiao, Du, Chao, Pang, Tianyu, Yang, Zhuoran, Hong, Mingyi, Tan, Vincent Y. F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Note on the Convergence of Muon
by: Li, Jiaxiang, et al.
Published: (2025)
by: Li, Jiaxiang, et al.
Published: (2025)
Parameter-free Algorithms for the Stochastically Extended Adversarial Model
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
by: Kunstner, Frederik, et al.
Published: (2024)
by: Kunstner, Frederik, et al.
Published: (2024)
A Correspondence-Driven Approach for Bilevel Decision-making with Nonconvex Lower-Level Problems
by: Jiang, Xiaotian, et al.
Published: (2025)
by: Jiang, Xiaotian, et al.
Published: (2025)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
by: Li, Binghui, et al.
Published: (2026)
by: Li, Binghui, et al.
Published: (2026)
A Barrier Function Approach for Bilevel Optimization with Coupled Lower-Level Constraints: Formulation, Approximation and Algorithms
by: Jiang, Xiaotian, et al.
Published: (2024)
by: Jiang, Xiaotian, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Problem-Parameter-Free Decentralized Nonconvex Stochastic Optimization
by: Li, Jiaxiang, et al.
Published: (2024)
by: Li, Jiaxiang, et al.
Published: (2024)
First-Order Algorithms Without Lipschitz Gradient: A Sequential Local Optimization Approach
by: Zhang, Junyu, et al.
Published: (2020)
by: Zhang, Junyu, et al.
Published: (2020)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)
by: Hou, Yunlong, et al.
Published: (2025)
An Efficient Memory Gradient Method for Extreme M-Eigenvalues of Elastic type Tensors
by: Du, Zhuolin, et al.
Published: (2026)
by: Du, Zhuolin, et al.
Published: (2026)
End-to-End Learning of Correlated Operating Reserve Requirements in Security-Constrained Economic Dispatch
by: Shen, Owen, et al.
Published: (2026)
by: Shen, Owen, et al.
Published: (2026)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Distributionally Robust Control with End-to-End Statistically Guaranteed Metric Learning
by: Wu, Jingyi, et al.
Published: (2025)
by: Wu, Jingyi, et al.
Published: (2025)
Optimization Outperforms Unscented Techniques for Nonlinear Smoothing
by: Howell, Payton, et al.
Published: (2025)
by: Howell, Payton, et al.
Published: (2025)
A Discretization Approach for Bilevel Optimization with Low-Dimensional and Non-Convex Lower-Level
by: Jiang, Xiaotian, et al.
Published: (2025)
by: Jiang, Xiaotian, et al.
Published: (2025)
On the Nature of Regularity Assumptions in Bilevel Optimization with Constrained Lower-level Problem
by: Jiang, Xiaotian, et al.
Published: (2026)
by: Jiang, Xiaotian, et al.
Published: (2026)
Distributed Stochastic Optimization under Heavy-Tailed Noises
by: Sun, Chao, et al.
Published: (2023)
by: Sun, Chao, et al.
Published: (2023)
Understanding the Variance Dichotomy in Continuous Simulation Optimization: A Minimax Lower Bound Perspective
by: Du, Jianzhong, et al.
Published: (2026)
by: Du, Jianzhong, et al.
Published: (2026)
Demystifying the Slash Pattern in Attention: The Role of RoPE
by: Cheng, Yuan, et al.
Published: (2026)
by: Cheng, Yuan, et al.
Published: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
by: Hong, Yusu, et al.
Published: (2024)
by: Hong, Yusu, et al.
Published: (2024)
Median-of-Means for Nash Equilibrium Seeking in Heavy-Tailed Games
by: Sun, Chao, et al.
Published: (2026)
by: Sun, Chao, et al.
Published: (2026)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
A Doubly Stochastically Perturbed Algorithm for Linearly Constrained Bilevel Optimization
by: Khanduri, Prashant, et al.
Published: (2025)
by: Khanduri, Prashant, et al.
Published: (2025)
Distributed Generalized Nash Equilibria Learning for Online Stochastic Aggregative Games
by: Du, Kaixin, et al.
Published: (2025)
by: Du, Kaixin, et al.
Published: (2025)
Decentralized Learning with Dynamically Refined Edge Weights: A Data-Dependent Framework
by: Du, Rongxing, et al.
Published: (2026)
by: Du, Rongxing, et al.
Published: (2026)
Data-Driven Distributionally Robust Mixed-Integer Control through Lifted Control Policy
by: Ma, Xutao, et al.
Published: (2025)
by: Ma, Xutao, et al.
Published: (2025)
LiMuon: Light and Fast Muon Optimizer for Large Models
by: Huang, Feihu, et al.
Published: (2025)
by: Huang, Feihu, et al.
Published: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
by: Yang, Penghui, et al.
Published: (2025)
by: Yang, Penghui, et al.
Published: (2025)
LEARN: An Invex Loss for Outlier Oblivious Robust Online Optimization
by: Barik, Adarsh, et al.
Published: (2024)
by: Barik, Adarsh, et al.
Published: (2024)
Distributed Stochastic Optimization for Non-Smooth and Weakly Convex Problems under Heavy-Tailed Noise
by: Hu, Jun, et al.
Published: (2025)
by: Hu, Jun, et al.
Published: (2025)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Constrained Optimization with Compressed Gradients: A Dynamical Systems Perspective
by: Xia, Zhaoyue, et al.
Published: (2024)
by: Xia, Zhaoyue, et al.
Published: (2024)
Similar Items
-
A Note on the Convergence of Muon
by: Li, Jiaxiang, et al.
Published: (2025) -
Parameter-free Algorithms for the Stochastically Extended Adversarial Model
by: Wang, Shuche, et al.
Published: (2025) -
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025) -
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025) -
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
by: Kunstner, Frederik, et al.
Published: (2024)