An Exploration of Non-Euclidean Gradient Descent: Muon and its Many Variants
Fuente:
arXiv
Saved in:
| Main Authors: | Crawshaw, Michael, Modi, Chirag, Liu, Mingrui, Gower, Robert M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
by: Crawshaw, Michael, et al.
Published: (2026)
by: Crawshaw, Michael, et al.
Published: (2026)
Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026)
by: Parshakova, Tetiana, et al.
Published: (2026)
Federated Learning under Periodic Client Participation and Heterogeneous Data: A New Communication-Efficient Algorithm and Analysis
by: Crawshaw, Michael, et al.
Published: (2024)
by: Crawshaw, Michael, et al.
Published: (2024)
Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
Local Steps Speed Up Local GD for Heterogeneous Distributed Logistic Regression
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
ATLAS: Adapting Trajectory Lengths and Step-Size for Hamiltonian Monte Carlo
by: Modi, Chirag
Published: (2024)
by: Modi, Chirag
Published: (2024)
EigenVI: score-based variational inference with orthogonal function expansions
by: Cai, Diana, et al.
Published: (2024)
by: Cai, Diana, et al.
Published: (2024)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Euclidean Distance Matrix Completion via Asymmetric Projected Gradient Descent
by: Li, Yicheng, et al.
Published: (2025)
by: Li, Yicheng, et al.
Published: (2025)
A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
by: Xie, Shuo, et al.
Published: (2025)
by: Xie, Shuo, et al.
Published: (2025)
Batch and match: black-box variational inference with a score-based divergence
by: Cai, Diana, et al.
Published: (2024)
by: Cai, Diana, et al.
Published: (2024)
Batch, match, and patch: low-rank approximations for score-based variational inference
by: Modi, Chirag, et al.
Published: (2024)
by: Modi, Chirag, et al.
Published: (2024)
MuonAll: Muon Variant for Efficient Finetuning of Large Language Models
by: Page, Saurabh, et al.
Published: (2025)
by: Page, Saurabh, et al.
Published: (2025)
In Search of Adam's Secret Sauce
by: Orvieto, Antonio, et al.
Published: (2025)
by: Orvieto, Antonio, et al.
Published: (2025)
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Occam Gradient Descent
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
Learning at the Speed of Physics: Equilibrium Propagation on Oscillator Ising Machines
by: Gower, Alex
Published: (2025)
by: Gower, Alex
Published: (2025)
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
Enhancing Policy Gradient with the Polyak Step-Size Adaption
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
Curl Descent: Non-Gradient Learning Dynamics with Sign-Diverse Plasticity
by: Ninou, Hugo, et al.
Published: (2025)
by: Ninou, Hugo, et al.
Published: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
A Bootstrap Perspective on Stochastic Gradient Descent
by: Lan, Hongjian, et al.
Published: (2025)
by: Lan, Hongjian, et al.
Published: (2025)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Stochastic Gradient Descent for Nonparametric Additive Regression
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
by: Cheng, Xiang, et al.
Published: (2023)
by: Cheng, Xiang, et al.
Published: (2023)
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
Non-Coherent Over-the-Air Decentralized Gradient Descent
by: Michelusi, Nicolo'
Published: (2022)
by: Michelusi, Nicolo'
Published: (2022)
Generative Modeling from Black-box Corruptions via Self-Consistent Stochastic Interpolants
by: Modi, Chirag, et al.
Published: (2025)
by: Modi, Chirag, et al.
Published: (2025)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
The Relative Gaussian Mechanism and its Application to Private Gradient Descent
by: Hendrikx, Hadrien, et al.
Published: (2023)
by: Hendrikx, Hadrien, et al.
Published: (2023)
Learning Curves of Stochastic Gradient Descent in Kernel Regression
by: Zhang, Haihan, et al.
Published: (2025)
by: Zhang, Haihan, et al.
Published: (2025)
Accelerated Gradient Descent for Faster Convergence with Minimal Overhead
by: Graca, Manuel, et al.
Published: (2026)
by: Graca, Manuel, et al.
Published: (2026)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Tight Generalization Error Bounds for Stochastic Gradient Descent in Non-convex Learning
by: Xiong, Wenjun, et al.
Published: (2025)
by: Xiong, Wenjun, et al.
Published: (2025)
Distributed Gradient Descent for Functional Learning
by: Yu, Zhan, et al.
Published: (2023)
by: Yu, Zhan, et al.
Published: (2023)
Similar Items
-
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026) -
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
by: Crawshaw, Michael, et al.
Published: (2026) -
Complexity Lower Bounds of Adaptive Gradient Algorithms for Non-convex Stochastic Optimization under Relaxed Smoothness
by: Crawshaw, Michael, et al.
Published: (2025) -
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026) -
Federated Learning under Periodic Client Participation and Heterogeneous Data: A New Communication-Efficient Algorithm and Analysis
by: Crawshaw, Michael, et al.
Published: (2024)