Drop-Muon: Update Less, Converge Faster
Fuente:
arXiv
Saved in:
| Main Authors: | Gruntkowska, Kaja, Maziane, Yassine, Qu, Zheng, Richtárik, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error Feedback for Muon and Friends
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
The Ball-Proximal (="Broximal") Point Method: a New Algorithm, Convergence Theory, and Applications
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle
by: Richtárik, Peter, et al.
Published: (2026)
by: Richtárik, Peter, et al.
Published: (2026)
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
by: Gruntkowska, Kaja, et al.
Published: (2024)
by: Gruntkowska, Kaja, et al.
Published: (2024)
Freya PAGE: First Optimal Time Complexity for Large-Scale Nonconvex Finite-Sum Optimization with Heterogeneous Asynchronous Computations
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Tighter Performance Theory of FedExProx
by: Anyszka, Wojciech, et al.
Published: (2024)
by: Anyszka, Wojciech, et al.
Published: (2024)
LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging
by: Maziane, Yassine, et al.
Published: (2026)
by: Maziane, Yassine, et al.
Published: (2026)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
Communication Compression for Byzantine Robust Learning: New Efficient Algorithms and Improved Rates
by: Rammal, Ahmad, et al.
Published: (2023)
by: Rammal, Ahmad, et al.
Published: (2023)
Beyond the Ideal: Analyzing the Inexact Muon Update
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Stabilized Proximal Point Method via Trust Region Control
by: Li, Hanmin, et al.
Published: (2026)
by: Li, Hanmin, et al.
Published: (2026)
Broximal Alignment for Global Non-Convex Optimization
by: Gruntkowska, Kaja, et al.
Published: (2026)
by: Gruntkowska, Kaja, et al.
Published: (2026)
Convergence Analysis of the PAGE Stochastic Algorithm for Weakly Convex Finite-Sum Optimization
by: Condat, Laurent, et al.
Published: (2025)
by: Condat, Laurent, et al.
Published: (2025)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
On the Convergence of DP-SGD with Adaptive Clipping
by: Shulgin, Egor, et al.
Published: (2024)
by: Shulgin, Egor, et al.
Published: (2024)
Convergence of Muon with Newton-Schulz
by: Kim, Gyu Yeol, et al.
Published: (2026)
by: Kim, Gyu Yeol, et al.
Published: (2026)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026)
by: Parshakova, Tetiana, et al.
Published: (2026)
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
A Computation and Communication Efficient Method for Distributed Nonconvex Problems in the Partial Participation Setting
by: Tyurin, Alexander, et al.
Published: (2022)
by: Tyurin, Alexander, et al.
Published: (2022)
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
by: Sokolov, Igor, et al.
Published: (2024)
by: Sokolov, Igor, et al.
Published: (2024)
On the Convergence Analysis of Muon
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
Faster Convergence of Local SGD for Over-Parameterized Models
by: Qin, Tiancheng, et al.
Published: (2022)
by: Qin, Tiancheng, et al.
Published: (2022)
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
A Unified Theory of Stochastic Proximal Point Methods without Smoothness
by: Richtárik, Peter, et al.
Published: (2024)
by: Richtárik, Peter, et al.
Published: (2024)
BiCoLoR: Communication-Efficient Optimization with Bidirectional Compression and Local Training
by: Condat, Laurent, et al.
Published: (2026)
by: Condat, Laurent, et al.
Published: (2026)
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2023)
by: Gorbunov, Eduard, et al.
Published: (2023)
Convergence for Discrete Parameter Update Schemes
by: Wilson, Paul, et al.
Published: (2025)
by: Wilson, Paul, et al.
Published: (2025)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Incremental Quasi-Newton Methods with Faster Superlinear Convergence Rates
by: Liu, Zhuanghua, et al.
Published: (2024)
by: Liu, Zhuanghua, et al.
Published: (2024)
Correlated Quantization for Faster Nonconvex Distributed Optimization
by: Panferov, Andrei, et al.
Published: (2024)
by: Panferov, Andrei, et al.
Published: (2024)
Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training
by: Southworth, Ben S., et al.
Published: (2026)
by: Southworth, Ben S., et al.
Published: (2026)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
Convergence of Distributed Adaptive Optimization with Local Updates
by: Cheng, Ziheng, et al.
Published: (2024)
by: Cheng, Ziheng, et al.
Published: (2024)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Similar Items
-
Error Feedback for Muon and Friends
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
by: Riabinin, Artem, et al.
Published: (2025) -
The Ball-Proximal (="Broximal") Point Method: a New Algorithm, Convergence Theory, and Applications
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle
by: Richtárik, Peter, et al.
Published: (2026)