MuonBP: Faster Muon via Block-Periodic Orthogonalization
Fuente:
arXiv
Saved in:
| Main Authors: | Khaled, Ahmed, Ozkara, Kaan, Yu, Tao, Hong, Mingyi, Park, Youngsuk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026)
by: Parshakova, Tetiana, et al.
Published: (2026)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
by: Ozkara, Kaan, et al.
Published: (2024)
by: Ozkara, Kaan, et al.
Published: (2024)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
LiMuon: Light and Fast Muon Optimizer for Large Models
by: Huang, Feihu, et al.
Published: (2025)
by: Huang, Feihu, et al.
Published: (2025)
A Note on the Convergence of Muon
by: Li, Jiaxiang, et al.
Published: (2025)
by: Li, Jiaxiang, et al.
Published: (2025)
Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training
by: Southworth, Ben S., et al.
Published: (2026)
by: Southworth, Ben S., et al.
Published: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Error Feedback for Muon and Friends
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Convergence of Muon with Newton-Schulz
by: Kim, Gyu Yeol, et al.
Published: (2026)
by: Kim, Gyu Yeol, et al.
Published: (2026)
Lions and Muons: Optimization via Stochastic Frank-Wolfe
by: Sfyraki, Maria-Eleni, et al.
Published: (2025)
by: Sfyraki, Maria-Eleni, et al.
Published: (2025)
Stochastic Rounding for LLM Training: Theory and Practice
by: Ozkara, Kaan, et al.
Published: (2025)
by: Ozkara, Kaan, et al.
Published: (2025)
Insights on Muon from Simple Quadratics
by: Gonon, Antoine, et al.
Published: (2026)
by: Gonon, Antoine, et al.
Published: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
On the Convergence Analysis of Muon
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Beyond the Ideal: Analyzing the Inexact Muon Update
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
by: He, Chuan, et al.
Published: (2025)
by: He, Chuan, et al.
Published: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
by: Li, Binghui, et al.
Published: (2026)
by: Li, Binghui, et al.
Published: (2026)
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
by: Takezawa, Yuki, et al.
Published: (2025)
by: Takezawa, Yuki, et al.
Published: (2025)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
Towards a Principled Muon under $μ\mathsf{P}$: Ensuring Spectral Conditions throughout Training
by: Zhao, John
Published: (2026)
by: Zhao, John
Published: (2026)
Faster Randomized Methods for Orthogonality Constrained Problems
by: Shustin, Boris, et al.
Published: (2021)
by: Shustin, Boris, et al.
Published: (2021)
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
by: Shumaylov, Zakhar, et al.
Published: (2026)
by: Shumaylov, Zakhar, et al.
Published: (2026)
Tuning-Free Stochastic Optimization
by: Khaled, Ahmed, et al.
Published: (2024)
by: Khaled, Ahmed, et al.
Published: (2024)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
by: Zhang, Minxin, et al.
Published: (2026)
by: Zhang, Minxin, et al.
Published: (2026)
DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent Method
by: Khaled, Ahmed, et al.
Published: (2023)
by: Khaled, Ahmed, et al.
Published: (2023)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Stochastic Approximation with Block Coordinate Optimal Stepsizes
by: Jiang, Tao, et al.
Published: (2025)
by: Jiang, Tao, et al.
Published: (2025)
Similar Items
-
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025) -
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026) -
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026) -
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)