Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nagashima, Shuntaro, Iiduka, Hideaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
von: Zhou, Dongruo, et al.
Veröffentlicht: (2018)
von: Zhou, Dongruo, et al.
Veröffentlicht: (2018)
Sharper Convergence Rates for Nonconvex Optimisation via Reduction Mappings
von: Markou, Evan, et al.
Veröffentlicht: (2025)
von: Markou, Evan, et al.
Veröffentlicht: (2025)
Deterministic Nonsmooth Nonconvex Optimization
von: Jordan, Michael I., et al.
Veröffentlicht: (2023)
von: Jordan, Michael I., et al.
Veröffentlicht: (2023)
Decentralized Sum-of-Nonconvex Optimization
von: Liu, Zhuanghua, et al.
Veröffentlicht: (2024)
von: Liu, Zhuanghua, et al.
Veröffentlicht: (2024)
Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double Optimism
von: Patitucci, Francisco, et al.
Veröffentlicht: (2025)
von: Patitucci, Francisco, et al.
Veröffentlicht: (2025)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
von: Wu, Haotian
Veröffentlicht: (2025)
von: Wu, Haotian
Veröffentlicht: (2025)
Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
LiMuon: Light and Fast Muon Optimizer for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2025)
von: Huang, Feihu, et al.
Veröffentlicht: (2025)
Convergence of Muon with Newton-Schulz
von: Kim, Gyu Yeol, et al.
Veröffentlicht: (2026)
von: Kim, Gyu Yeol, et al.
Veröffentlicht: (2026)
Convergence Bound and Critical Batch Size of Muon Optimizer
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
von: Sato, Naoki, et al.
Veröffentlicht: (2025)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Improved Sample Complexity for Private Nonsmooth Nonconvex Optimization
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
Online Nonconvex Bilevel Optimization with Bregman Divergences
von: Bohne, Jason, et al.
Veröffentlicht: (2024)
von: Bohne, Jason, et al.
Veröffentlicht: (2024)
Improved Learning Rates for Stochastic Optimization
von: Li, Shaojie, et al.
Veröffentlicht: (2021)
von: Li, Shaojie, et al.
Veröffentlicht: (2021)
Parametric Nonconvex Optimization via Convex Surrogates
von: Wang, Renzi, et al.
Veröffentlicht: (2026)
von: Wang, Renzi, et al.
Veröffentlicht: (2026)
Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization
von: Gong, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Gong, Xiaochuan, et al.
Veröffentlicht: (2025)
Muon Optimizes Under Spectral Norm Constraints
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2024)
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2024)
The Newton-Muon Optimizer
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization
von: Xiao, Nachuan, et al.
Veröffentlicht: (2024)
von: Xiao, Nachuan, et al.
Veröffentlicht: (2024)
Decentralized Stochastic Nonconvex Optimization under the Relaxed Smoothness
von: Luo, Luo, et al.
Veröffentlicht: (2025)
von: Luo, Luo, et al.
Veröffentlicht: (2025)
On the Complexity of Decentralized Smooth Nonconvex Finite-Sum Optimization
von: Luo, Luo, et al.
Veröffentlicht: (2022)
von: Luo, Luo, et al.
Veröffentlicht: (2022)
On the Hardness of Meaningful Local Guarantees in Nonsmooth Nonconvex Optimization
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
Muon Does Not Converge on Convex Lipschitz Functions
von: Parshakova, Tetiana, et al.
Veröffentlicht: (2026)
von: Parshakova, Tetiana, et al.
Veröffentlicht: (2026)
Drop-Muon: Update Less, Converge Faster
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
Lions and Muons: Optimization via Stochastic Frank-Wolfe
von: Sfyraki, Maria-Eleni, et al.
Veröffentlicht: (2025)
von: Sfyraki, Maria-Eleni, et al.
Veröffentlicht: (2025)
Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization
von: Qin, Zhen, et al.
Veröffentlicht: (2023)
von: Qin, Zhen, et al.
Veröffentlicht: (2023)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
On the Convergence Analysis of Muon
von: Shen, Wei, et al.
Veröffentlicht: (2025)
von: Shen, Wei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026) -
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023) -
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023) -
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024) -
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)