Preconditioning Benefits of Spectral Orthogonalization in Muon
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Jianhao, Huang, Yu, Chi, Yuejie, Chen, Yuxin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025)
by: Chi, Yuejie, et al.
Published: (2025)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
by: Ma, Jianhao, et al.
Published: (2025)
by: Ma, Jianhao, et al.
Published: (2025)
The Power of Preconditioning in Overparameterized Low-Rank Matrix Sensing
by: Xu, Xingyu, et al.
Published: (2023)
by: Xu, Xingyu, et al.
Published: (2023)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Tree-Preconditioned Differentiable Optimization and Axioms as Layers
by: Liao, Yuexin
Published: (2025)
by: Liao, Yuexin
Published: (2025)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
by: He, Chuan, et al.
Published: (2025)
by: He, Chuan, et al.
Published: (2025)
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Optimal Control Operator Perspective and a Neural Adaptive Spectral Method
by: Feng, Mingquan, et al.
Published: (2024)
by: Feng, Mingquan, et al.
Published: (2024)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
by: Ma, Shaocong, et al.
Published: (2026)
by: Ma, Shaocong, et al.
Published: (2026)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
Primal-Dual Spectral Representation for Off-policy Evaluation
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
The Sample-Communication Complexity Trade-off in Federated Q-Learning
by: Salgia, Sudeep, et al.
Published: (2024)
by: Salgia, Sudeep, et al.
Published: (2024)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
by: Shumaylov, Zakhar, et al.
Published: (2026)
by: Shumaylov, Zakhar, et al.
Published: (2026)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
by: Yan, Mingsong, et al.
Published: (2026)
by: Yan, Mingsong, et al.
Published: (2026)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Robust Reinforcement Learning in Finance: Modeling Market Impact with Elliptic Uncertainty Sets
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
by: Ma, Shaocong, et al.
Published: (2025)
by: Ma, Shaocong, et al.
Published: (2025)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
LiMuon: Light and Fast Muon Optimizer for Large Models
by: Huang, Feihu, et al.
Published: (2025)
by: Huang, Feihu, et al.
Published: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Towards Stable Machine Learning Model Retraining via Slowly Varying Sequences
by: Bertsimas, Dimitris, et al.
Published: (2024)
by: Bertsimas, Dimitris, et al.
Published: (2024)
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2025)
by: Huang, Yilie, et al.
Published: (2025)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
by: Oikonomidis, Konstantinos, et al.
Published: (2026)
by: Oikonomidis, Konstantinos, et al.
Published: (2026)
Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
by: Li, Gen, et al.
Published: (2020)
by: Li, Gen, et al.
Published: (2020)
Sublinear Regret for a Class of Continuous-Time Linear-Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2024)
by: Huang, Yilie, et al.
Published: (2024)
Similar Items
-
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025) -
Statistical and Algorithmic Foundations of Reinforcement Learning
by: Chi, Yuejie, et al.
Published: (2025) -
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026) -
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026) -
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)