On the Convergence Analysis of Muon
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Wei, Huang, Ruichuan, Huang, Minhui, Shen, Cong, Zhang, Jiawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
di: Shen, Wei, et al.
Pubblicazione: (2025)
di: Shen, Wei, et al.
Pubblicazione: (2025)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
di: Shen, Wei, et al.
Pubblicazione: (2023)
di: Shen, Wei, et al.
Pubblicazione: (2023)
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
di: Shen, Wei, et al.
Pubblicazione: (2025)
di: Shen, Wei, et al.
Pubblicazione: (2025)
Stochastic Smoothed Primal-Dual Algorithms for Nonconvex Optimization with Linear Inequality Constraints
di: Huang, Ruichuan, et al.
Pubblicazione: (2025)
di: Huang, Ruichuan, et al.
Pubblicazione: (2025)
Group Projected Subspace Pursuit for Block Sparse Signal Reconstruction: Convergence Analysis and Applications
di: He, Roy Y., et al.
Pubblicazione: (2024)
di: He, Roy Y., et al.
Pubblicazione: (2024)
Accelerating Convergence of Score-Based Diffusion Models, Provably
di: Li, Gen, et al.
Pubblicazione: (2024)
di: Li, Gen, et al.
Pubblicazione: (2024)
Convergence of linear programming hierarchies for Gibbs states of spin systems
di: Fawzi, Hamza, et al.
Pubblicazione: (2025)
di: Fawzi, Hamza, et al.
Pubblicazione: (2025)
On the Robustness of Cross-Concentrated Sampling for Matrix Completion
di: Cai, HanQin, et al.
Pubblicazione: (2024)
di: Cai, HanQin, et al.
Pubblicazione: (2024)
High-probability sample complexities for policy evaluation with linear function approximation
di: Li, Gen, et al.
Pubblicazione: (2023)
di: Li, Gen, et al.
Pubblicazione: (2023)
Stochastic Zeroth-Order Optimization under Strongly Convexity and Lipschitz Hessian: Minimax Sample Complexity
di: Yu, Qian, et al.
Pubblicazione: (2024)
di: Yu, Qian, et al.
Pubblicazione: (2024)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
di: Li, Gen, et al.
Pubblicazione: (2021)
di: Li, Gen, et al.
Pubblicazione: (2021)
Convergence of Muon with Newton-Schulz
di: Kim, Gyu Yeol, et al.
Pubblicazione: (2026)
di: Kim, Gyu Yeol, et al.
Pubblicazione: (2026)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
Wasserstein Distributionally Robust Estimation in High Dimensions: Performance Analysis and Optimal Hyperparameter Tuning
di: Aolaritei, Liviu, et al.
Pubblicazione: (2022)
di: Aolaritei, Liviu, et al.
Pubblicazione: (2022)
Adversarial Water-Filling: Theory, Algorithms and Foundation Model
di: Tong, Xindi, et al.
Pubblicazione: (2026)
di: Tong, Xindi, et al.
Pubblicazione: (2026)
Escaping Saddle Points for Nonsmooth Weakly Convex Functions via Perturbed Proximal Algorithms
di: Huang, Minhui, et al.
Pubblicazione: (2021)
di: Huang, Minhui, et al.
Pubblicazione: (2021)
LiMuon: Light and Fast Muon Optimizer for Large Models
di: Huang, Feihu, et al.
Pubblicazione: (2025)
di: Huang, Feihu, et al.
Pubblicazione: (2025)
Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGD
di: Wang, Jiayi, et al.
Pubblicazione: (2020)
di: Wang, Jiayi, et al.
Pubblicazione: (2020)
Drop-Muon: Update Less, Converge Faster
di: Gruntkowska, Kaja, et al.
Pubblicazione: (2025)
di: Gruntkowska, Kaja, et al.
Pubblicazione: (2025)
Muon Does Not Converge on Convex Lipschitz Functions
di: Parshakova, Tetiana, et al.
Pubblicazione: (2026)
di: Parshakova, Tetiana, et al.
Pubblicazione: (2026)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
di: Li, Gen, et al.
Pubblicazione: (2020)
di: Li, Gen, et al.
Pubblicazione: (2020)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
di: Yang, Tong, et al.
Pubblicazione: (2025)
di: Yang, Tong, et al.
Pubblicazione: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
di: Yang, Tong, et al.
Pubblicazione: (2024)
di: Yang, Tong, et al.
Pubblicazione: (2024)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
di: Nagashima, Shuntaro, et al.
Pubblicazione: (2026)
di: Nagashima, Shuntaro, et al.
Pubblicazione: (2026)
Linear regression with overparameterized linear neural networks: Tight upper and lower bounds for implicit $\ell^1$-regularization
di: Matt, Hannes, et al.
Pubblicazione: (2025)
di: Matt, Hannes, et al.
Pubblicazione: (2025)
Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control
di: Liu, Yujie, et al.
Pubblicazione: (2025)
di: Liu, Yujie, et al.
Pubblicazione: (2025)
A Neural Network Algorithm for KL Divergence Estimation with Quantitative Error Bounds
di: Foss, Mikil, et al.
Pubblicazione: (2025)
di: Foss, Mikil, et al.
Pubblicazione: (2025)
A Dual Basis Approach for Structured Robust Euclidean Distance Geometry
di: Kundu, Chandra, et al.
Pubblicazione: (2025)
di: Kundu, Chandra, et al.
Pubblicazione: (2025)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
di: Zurek, Matthew, et al.
Pubblicazione: (2025)
Convexity in Disguise: A Theoretical Framework for Nonconvex Low-Rank Matrix Estimation
di: Cui, Chengyu, et al.
Pubblicazione: (2026)
di: Cui, Chengyu, et al.
Pubblicazione: (2026)
Recovering Simultaneously Structured Data via Non-Convex Iteratively Reweighted Least Squares
di: Kümmerle, Christian, et al.
Pubblicazione: (2023)
di: Kümmerle, Christian, et al.
Pubblicazione: (2023)
Generalized Orthogonal Procrustes Problem under Arbitrary Adversaries
di: Ling, Shuyang
Pubblicazione: (2021)
di: Ling, Shuyang
Pubblicazione: (2021)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
di: Zurek, Matthew, et al.
Pubblicazione: (2024)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
di: Zamir, Guy, et al.
Pubblicazione: (2026)
di: Zamir, Guy, et al.
Pubblicazione: (2026)
Tight Regret Bounds for Bayesian Optimization in One Dimension
di: Scarlett, Jonathan
Pubblicazione: (2018)
di: Scarlett, Jonathan
Pubblicazione: (2018)
Variational Inference on the Boolean Hypercube with the Quantum Entropy
di: Beyler, Eliot, et al.
Pubblicazione: (2024)
di: Beyler, Eliot, et al.
Pubblicazione: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
di: Zurek, Matthew, et al.
Pubblicazione: (2023)
di: Zurek, Matthew, et al.
Pubblicazione: (2023)
Structured Sampling for Robust Euclidean Distance Geometry
di: Kundu, Chandra, et al.
Pubblicazione: (2024)
di: Kundu, Chandra, et al.
Pubblicazione: (2024)
The augmented NLP bound for maximum-entropy remote sampling
di: Ponte, Gabriel, et al.
Pubblicazione: (2026)
di: Ponte, Gabriel, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
di: Shen, Wei, et al.
Pubblicazione: (2025) -
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
di: Shen, Wei, et al.
Pubblicazione: (2023) -
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
di: Shen, Wei, et al.
Pubblicazione: (2025) -
Stochastic Smoothed Primal-Dual Algorithms for Nonconvex Optimization with Linear Inequality Constraints
di: Huang, Ruichuan, et al.
Pubblicazione: (2025) -
Group Projected Subspace Pursuit for Block Sparse Signal Reconstruction: Convergence Analysis and Applications
di: He, Roy Y., et al.
Pubblicazione: (2024)