Muon Does Not Converge on Convex Lipschitz Functions
Fuente:
arXiv
Saved in:
| Main Authors: | Parshakova, Tetiana, Khaled, Ahmed, Crawshaw, Michael, Garrigos, Guillaume, Gower, Robert M. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Handbook of Convergence Theorems for (Stochastic) Gradient Methods
by: Garrigos, Guillaume, et al.
Published: (2023)
by: Garrigos, Guillaume, et al.
Published: (2023)
Online Inventory Problems: Beyond the i.i.d. Setting with Online Convex Optimization
by: Hihat, Massil, et al.
Published: (2023)
by: Hihat, Massil, et al.
Published: (2023)
Factor Fitting, Rank Allocation, and Partitioning in Multilevel Low Rank Matrices
by: Parshakova, Tetiana, et al.
Published: (2023)
by: Parshakova, Tetiana, et al.
Published: (2023)
Tracking the Median of Gradients with a Stochastic Proximal Point Method
by: Schaipp, Fabian, et al.
Published: (2024)
by: Schaipp, Fabian, et al.
Published: (2024)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
Improving Convergence and Generalization Using Parameter Symmetries
by: Zhao, Bo, et al.
Published: (2023)
by: Zhao, Bo, et al.
Published: (2023)
Convergence of Muon with Newton-Schulz
by: Kim, Gyu Yeol, et al.
Published: (2026)
by: Kim, Gyu Yeol, et al.
Published: (2026)
Bias-Optimal Bounds for SGD: A Computer-Aided Lyapunov Analysis
by: Cortild, Daniel, et al.
Published: (2025)
by: Cortild, Daniel, et al.
Published: (2025)
Multivariate Online Linear Regression for Hierarchical Forecasting
by: Hihat, Massil, et al.
Published: (2024)
by: Hihat, Massil, et al.
Published: (2024)
On Convergence of Incremental Gradient for Non-Convex Smooth Functions
by: Koloskova, Anastasia, et al.
Published: (2023)
by: Koloskova, Anastasia, et al.
Published: (2023)
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
On the Convergence Analysis of Muon
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Stochastic Weakly Convex Optimization Beyond Lipschitz Continuity
by: Gao, Wenzhi, et al.
Published: (2024)
by: Gao, Wenzhi, et al.
Published: (2024)
Level Set Teleportation: An Optimization Perspective
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
Revisiting Subgradient Method: Complexity and Convergence Beyond Lipschitz Continuity
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Convergence Analysis of the Wasserstein Proximal Algorithm beyond Geodesic Convexity
by: Zhu, Shuailong, et al.
Published: (2025)
by: Zhu, Shuailong, et al.
Published: (2025)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
Tuning-Free Stochastic Optimization
by: Khaled, Ahmed, et al.
Published: (2024)
by: Khaled, Ahmed, et al.
Published: (2024)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
LiMuon: Light and Fast Muon Optimizer for Large Models
by: Huang, Feihu, et al.
Published: (2025)
by: Huang, Feihu, et al.
Published: (2025)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
Improved Last-Iterate Convergence of Shuffling Gradient Methods for Nonsmooth Convex Optimization
by: Liu, Zijian, et al.
Published: (2025)
by: Liu, Zijian, et al.
Published: (2025)
Convergence Analysis of the PAGE Stochastic Algorithm for Weakly Convex Finite-Sum Optimization
by: Condat, Laurent, et al.
Published: (2025)
by: Condat, Laurent, et al.
Published: (2025)
Optimization Algorithm Design via Electric Circuits
by: Boyd, Stephen P., et al.
Published: (2024)
by: Boyd, Stephen P., et al.
Published: (2024)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Error Feedback for Muon and Friends
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
by: Khaled, Ahmed, et al.
Published: (2023)
by: Khaled, Ahmed, et al.
Published: (2023)
Last-Iterate Complexity of SGD for Convex and Smooth Stochastic Problems
by: Garrigos, Guillaume, et al.
Published: (2025)
by: Garrigos, Guillaume, et al.
Published: (2025)
Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
by: Schaipp, Fabian, et al.
Published: (2026)
by: Schaipp, Fabian, et al.
Published: (2026)
Insights on Muon from Simple Quadratics
by: Gonon, Antoine, et al.
Published: (2026)
by: Gonon, Antoine, et al.
Published: (2026)
A Theoretical and Empirical Study on the Convergence of Adam with an "Exact" Constant Step Size in Non-Convex Settings
by: Mazumder, Alokendu, et al.
Published: (2023)
by: Mazumder, Alokendu, et al.
Published: (2023)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Quantitative Convergence Analysis of Projected Stochastic Gradient Descent for Non-Convex Losses via the Goldstein Subdifferential
by: Zheng, Yuping, et al.
Published: (2025)
by: Zheng, Yuping, et al.
Published: (2025)
A Randomized Linearly Convergent Frank-Wolfe-type Method for Smooth Convex Minimization over the Spectrahedron
by: Garber, Dan
Published: (2025)
by: Garber, Dan
Published: (2025)
Similar Items
-
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026) -
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024) -
Handbook of Convergence Theorems for (Stochastic) Gradient Methods
by: Garrigos, Guillaume, et al.
Published: (2023) -
Online Inventory Problems: Beyond the i.i.d. Setting with Online Convex Optimization
by: Hihat, Massil, et al.
Published: (2023) -
Factor Fitting, Rank Allocation, and Partitioning in Multilevel Low Rank Matrices
by: Parshakova, Tetiana, et al.
Published: (2023)