Phases of Muon: When Muon Eclipses SignSGD
Fuente:
arXiv
Guardado en:
| Autores principales: | Paquette, Elliot, Marshall, Noah, Benigni, Lucas, Wang, Guangyuan, Agarwala, Atish, Paquette, Courtney |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dimension-adapted Momentum Outscales SGD
por: Ferbach, Damien, et al.
Publicado: (2025)
por: Ferbach, Damien, et al.
Publicado: (2025)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
por: Petrov, Egor, et al.
Publicado: (2025)
por: Petrov, Egor, et al.
Publicado: (2025)
4+3 Phases of Compute-Optimal Neural Scaling Laws
por: Paquette, Elliot, et al.
Publicado: (2024)
por: Paquette, Elliot, et al.
Publicado: (2024)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
por: Marshall, Noah, et al.
Publicado: (2024)
por: Marshall, Noah, et al.
Publicado: (2024)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
por: Xiao, Ke Liang, et al.
Publicado: (2024)
por: Xiao, Ke Liang, et al.
Publicado: (2024)
Logarithmic-time Schedules for Scaling Language Models with Momentum
por: Ferbach, Damien, et al.
Publicado: (2026)
por: Ferbach, Damien, et al.
Publicado: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
por: Kim, Jihwan, et al.
Publicado: (2026)
por: Kim, Jihwan, et al.
Publicado: (2026)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
por: Everett, Katie, et al.
Publicado: (2026)
por: Everett, Katie, et al.
Publicado: (2026)
High-dimensional Limit of SGD for Diagonal Linear Networks
por: Malaxechebarría, Begoña García, et al.
Publicado: (2026)
por: Malaxechebarría, Begoña García, et al.
Publicado: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
por: Yu, Dingzhi, et al.
Publicado: (2026)
por: Yu, Dingzhi, et al.
Publicado: (2026)
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
por: Collins-Woodfin, Elizabeth, et al.
Publicado: (2024)
por: Collins-Woodfin, Elizabeth, et al.
Publicado: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
por: Tao, Hongyi, et al.
Publicado: (2026)
por: Tao, Hongyi, et al.
Publicado: (2026)
Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
por: Benigni, Lucas, et al.
Publicado: (2025)
por: Benigni, Lucas, et al.
Publicado: (2025)
Mirror Descent Algorithms with Nearly Dimension-Independent Rates for Differentially-Private Stochastic Saddle-Point Problems
por: González, Tomás, et al.
Publicado: (2024)
por: González, Tomás, et al.
Publicado: (2024)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
LiMuon: Light and Fast Muon Optimizer for Large Models
por: Huang, Feihu, et al.
Publicado: (2025)
por: Huang, Feihu, et al.
Publicado: (2025)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
por: Roulet, Vincent, et al.
Publicado: (2023)
por: Roulet, Vincent, et al.
Publicado: (2023)
Convergence of Muon with Newton-Schulz
por: Kim, Gyu Yeol, et al.
Publicado: (2026)
por: Kim, Gyu Yeol, et al.
Publicado: (2026)
Error Feedback for Muon and Friends
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
Insights on Muon from Simple Quadratics
por: Gonon, Antoine, et al.
Publicado: (2026)
por: Gonon, Antoine, et al.
Publicado: (2026)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
por: Agarwala, Atish, et al.
Publicado: (2024)
por: Agarwala, Atish, et al.
Publicado: (2024)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
por: Huang, Feihu, et al.
Publicado: (2026)
por: Huang, Feihu, et al.
Publicado: (2026)
Lions and Muons: Optimization via Stochastic Frank-Wolfe
por: Sfyraki, Maria-Eleni, et al.
Publicado: (2025)
por: Sfyraki, Maria-Eleni, et al.
Publicado: (2025)
The Newton-Muon Optimizer
por: Du, Zhehang, et al.
Publicado: (2026)
por: Du, Zhehang, et al.
Publicado: (2026)
On the Convergence Analysis of Muon
por: Shen, Wei, et al.
Publicado: (2025)
por: Shen, Wei, et al.
Publicado: (2025)
Muon Does Not Converge on Convex Lipschitz Functions
por: Parshakova, Tetiana, et al.
Publicado: (2026)
por: Parshakova, Tetiana, et al.
Publicado: (2026)
Muon is Provably Faster with Momentum Variance Reduction
por: Qian, Xun, et al.
Publicado: (2025)
por: Qian, Xun, et al.
Publicado: (2025)
Drop-Muon: Update Less, Converge Faster
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
Beyond the Ideal: Analyzing the Inexact Muon Update
por: Shulgin, Egor, et al.
Publicado: (2025)
por: Shulgin, Egor, et al.
Publicado: (2025)
Muon Optimizes Under Spectral Norm Constraints
por: Chen, Lizhang, et al.
Publicado: (2025)
por: Chen, Lizhang, et al.
Publicado: (2025)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
por: Li, Binghui, et al.
Publicado: (2026)
por: Li, Binghui, et al.
Publicado: (2026)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
por: Nagashima, Shuntaro, et al.
Publicado: (2026)
por: Nagashima, Shuntaro, et al.
Publicado: (2026)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
por: He, Chuan, et al.
Publicado: (2025)
por: He, Chuan, et al.
Publicado: (2025)
Sign-SGD via Parameter-Free Optimization
por: Medyakov, Daniil, et al.
Publicado: (2025)
por: Medyakov, Daniil, et al.
Publicado: (2025)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
por: Zhang, Minxin, et al.
Publicado: (2025)
por: Zhang, Minxin, et al.
Publicado: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
por: Fan, Chen, et al.
Publicado: (2025)
por: Fan, Chen, et al.
Publicado: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
por: Ma, Jianhao, et al.
Publicado: (2026)
por: Ma, Jianhao, et al.
Publicado: (2026)
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
por: Takezawa, Yuki, et al.
Publicado: (2025)
por: Takezawa, Yuki, et al.
Publicado: (2025)
Muon Dynamics as a Spectral Wasserstein Flow
por: Peyré, Gabriel
Publicado: (2026)
por: Peyré, Gabriel
Publicado: (2026)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
por: Choudhury, Sayantan, et al.
Publicado: (2026)
por: Choudhury, Sayantan, et al.
Publicado: (2026)
Ejemplares similares
-
Dimension-adapted Momentum Outscales SGD
por: Ferbach, Damien, et al.
Publicado: (2025) -
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
por: Petrov, Egor, et al.
Publicado: (2025) -
4+3 Phases of Compute-Optimal Neural Scaling Laws
por: Paquette, Elliot, et al.
Publicado: (2024) -
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
por: Marshall, Noah, et al.
Publicado: (2024) -
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
por: Xiao, Ke Liang, et al.
Publicado: (2024)