Insights on Muon from Simple Quadratics
Fuente:
arXiv
Guardado en:
| Autores principales: | Gonon, Antoine, Muşat, Andreea-Alexandra, Boumal, Nicolas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A non-autonomous center-stable set theorem for saddle avoidance in optimization
por: Muşat, Andreea-Alexandra, et al.
Publicado: (2026)
por: Muşat, Andreea-Alexandra, et al.
Publicado: (2026)
Gradient descent avoids strict saddles with a simple line-search method too
por: Muşat, Andreea-Alexandra, et al.
Publicado: (2025)
por: Muşat, Andreea-Alexandra, et al.
Publicado: (2025)
Distributional Adversarial Attacks and Training in Deep Hedging
por: He, Guangyi, et al.
Publicado: (2025)
por: He, Guangyi, et al.
Publicado: (2025)
Synchronization on circles and spheres with nonlinear interactions
por: Criscitiello, Christopher, et al.
Publicado: (2024)
por: Criscitiello, Christopher, et al.
Publicado: (2024)
Gradient Regularized Newton Boosting Trees with Global Convergence
por: Zozoulenko, Nikita, et al.
Publicado: (2026)
por: Zozoulenko, Nikita, et al.
Publicado: (2026)
Phases of Muon: When Muon Eclipses SignSGD
por: Paquette, Elliot, et al.
Publicado: (2026)
por: Paquette, Elliot, et al.
Publicado: (2026)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
por: Khaled, Ahmed, et al.
Publicado: (2025)
por: Khaled, Ahmed, et al.
Publicado: (2025)
LiMuon: Light and Fast Muon Optimizer for Large Models
por: Huang, Feihu, et al.
Publicado: (2025)
por: Huang, Feihu, et al.
Publicado: (2025)
Convergence of Muon with Newton-Schulz
por: Kim, Gyu Yeol, et al.
Publicado: (2026)
por: Kim, Gyu Yeol, et al.
Publicado: (2026)
Error Feedback for Muon and Friends
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
Gaussian Match-and-Copy: A Minimalist Benchmark for Studying Transformer Induction
por: Gonon, Antoine, et al.
Publicado: (2026)
por: Gonon, Antoine, et al.
Publicado: (2026)
Muon Does Not Converge on Convex Lipschitz Functions
por: Parshakova, Tetiana, et al.
Publicado: (2026)
por: Parshakova, Tetiana, et al.
Publicado: (2026)
Muon is Provably Faster with Momentum Variance Reduction
por: Qian, Xun, et al.
Publicado: (2025)
por: Qian, Xun, et al.
Publicado: (2025)
Drop-Muon: Update Less, Converge Faster
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
por: Gruntkowska, Kaja, et al.
Publicado: (2025)
Beyond the Ideal: Analyzing the Inexact Muon Update
por: Shulgin, Egor, et al.
Publicado: (2025)
por: Shulgin, Egor, et al.
Publicado: (2025)
Muon Optimizes Under Spectral Norm Constraints
por: Chen, Lizhang, et al.
Publicado: (2025)
por: Chen, Lizhang, et al.
Publicado: (2025)
Fast convergence to non-isolated minima: four equivalent conditions for $\mathrm{C}^2$ functions
por: Rebjock, Quentin, et al.
Publicado: (2023)
por: Rebjock, Quentin, et al.
Publicado: (2023)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
por: Nagashima, Shuntaro, et al.
Publicado: (2026)
por: Nagashima, Shuntaro, et al.
Publicado: (2026)
Lions and Muons: Optimization via Stochastic Frank-Wolfe
por: Sfyraki, Maria-Eleni, et al.
Publicado: (2025)
por: Sfyraki, Maria-Eleni, et al.
Publicado: (2025)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
por: Li, Binghui, et al.
Publicado: (2026)
por: Li, Binghui, et al.
Publicado: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
por: Zhang, Minxin, et al.
Publicado: (2025)
por: Zhang, Minxin, et al.
Publicado: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
por: Fan, Chen, et al.
Publicado: (2025)
por: Fan, Chen, et al.
Publicado: (2025)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
por: Huang, Feihu, et al.
Publicado: (2026)
por: Huang, Feihu, et al.
Publicado: (2026)
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
por: Takezawa, Yuki, et al.
Publicado: (2025)
por: Takezawa, Yuki, et al.
Publicado: (2025)
The Newton-Muon Optimizer
por: Du, Zhehang, et al.
Publicado: (2026)
por: Du, Zhehang, et al.
Publicado: (2026)
On the Convergence Analysis of Muon
por: Shen, Wei, et al.
Publicado: (2025)
por: Shen, Wei, et al.
Publicado: (2025)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
por: Choudhury, Sayantan, et al.
Publicado: (2026)
por: Choudhury, Sayantan, et al.
Publicado: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
por: Petrov, Egor, et al.
Publicado: (2025)
por: Petrov, Egor, et al.
Publicado: (2025)
Expressivity of Quadratic Neural ODEs
por: Hanson, Joshua, et al.
Publicado: (2025)
por: Hanson, Joshua, et al.
Publicado: (2025)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
por: Iiduka, Hideaki
Publicado: (2026)
por: Iiduka, Hideaki
Publicado: (2026)
On the Effectiveness of the z-Transform Method in Quadratic Optimization
por: Bach, Francis
Publicado: (2025)
por: Bach, Francis
Publicado: (2025)
Tight Rates for Bandit Control Beyond Quadratics
por: Sun, Y. Jennifer, et al.
Publicado: (2024)
por: Sun, Y. Jennifer, et al.
Publicado: (2024)
Accelerated Optimization Landscape of Linear-Quadratic Regulator
por: Feng, Lechen, et al.
Publicado: (2023)
por: Feng, Lechen, et al.
Publicado: (2023)
Towards a Principled Muon under $μ\mathsf{P}$: Ensuring Spectral Conditions throughout Training
por: Zhao, John
Publicado: (2026)
por: Zhao, John
Publicado: (2026)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
por: Riabinin, Artem, et al.
Publicado: (2025)
por: Riabinin, Artem, et al.
Publicado: (2025)
Sub-optimality of the Separation Principle for Quadratic Control from Bilinear Observations
por: Sattar, Yahya, et al.
Publicado: (2025)
por: Sattar, Yahya, et al.
Publicado: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
por: Ma, Jianhao, et al.
Publicado: (2026)
por: Ma, Jianhao, et al.
Publicado: (2026)
Optimization over bounded-rank matrices through a desingularization enables joint global and local guarantees
por: Rebjock, Quentin, et al.
Publicado: (2024)
por: Rebjock, Quentin, et al.
Publicado: (2024)
Fast convergence of trust-regions for non-isolated minima via analysis of CG on indefinite matrices
por: Rebjock, Quentin, et al.
Publicado: (2023)
por: Rebjock, Quentin, et al.
Publicado: (2023)
Quadratic Objective Perturbation: Curvature-Based Differential Privacy
por: Cortild, Daniel, et al.
Publicado: (2026)
por: Cortild, Daniel, et al.
Publicado: (2026)
Ejemplares similares
-
A non-autonomous center-stable set theorem for saddle avoidance in optimization
por: Muşat, Andreea-Alexandra, et al.
Publicado: (2026) -
Gradient descent avoids strict saddles with a simple line-search method too
por: Muşat, Andreea-Alexandra, et al.
Publicado: (2025) -
Distributional Adversarial Attacks and Training in Deep Hedging
por: He, Guangyi, et al.
Publicado: (2025) -
Synchronization on circles and spheres with nonlinear interactions
por: Criscitiello, Christopher, et al.
Publicado: (2024) -
Gradient Regularized Newton Boosting Trees with Global Convergence
por: Zozoulenko, Nikita, et al.
Publicado: (2026)