From SGD to Muon: Adaptive Optimization via Schatten-p Norms
Fuente:
arXiv
Saved in:
| Main Authors: | Massena, Thomas, Friedrich, Corentin, Serrurier, Mathieu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Turbo-Muon: Accelerating Orthogonality-Based Optimization with Pre-Conditioning
by: Boissin, Thibaut, et al.
Published: (2025)
by: Boissin, Thibaut, et al.
Published: (2025)
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
by: Massena, Thomas, et al.
Published: (2025)
by: Massena, Thomas, et al.
Published: (2025)
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
by: Boissin, Thibaut, et al.
Published: (2025)
by: Boissin, Thibaut, et al.
Published: (2025)
Fast and Flexible Robustness Certificates for Semantic Segmentation
by: Massena, Thomas, et al.
Published: (2025)
by: Massena, Thomas, et al.
Published: (2025)
DP-SGD Without Clipping: The Lipschitz Neural Network Way
by: Bethune, Louis, et al.
Published: (2023)
by: Bethune, Louis, et al.
Published: (2023)
Generating Heterogeneous Multi-dimensional Data : A Comparative Study
by: Corbeau, Michael, et al.
Published: (2025)
by: Corbeau, Michael, et al.
Published: (2025)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
by: Wei, Chengkun, et al.
Published: (2025)
by: Wei, Chengkun, et al.
Published: (2025)
On the explainable properties of 1-Lipschitz Neural Networks: An Optimal Transport Perspective
by: Serrurier, Mathieu, et al.
Published: (2022)
by: Serrurier, Mathieu, et al.
Published: (2022)
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
by: Cheng, Peng, et al.
Published: (2026)
by: Cheng, Peng, et al.
Published: (2026)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Muon Optimizer Accelerates Grokking
by: Tveit, Amund, et al.
Published: (2025)
by: Tveit, Amund, et al.
Published: (2025)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
by: Li, Yibang, et al.
Published: (2026)
by: Li, Yibang, et al.
Published: (2026)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Comparative Analysis of Novel NIRMAL Optimizer Against Adam and SGD with Momentum
by: Gaud, Nirmal, et al.
Published: (2025)
by: Gaud, Nirmal, et al.
Published: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
DynMuon: A Dynamic Spectral Shaping View of Muon
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
by: He, Chuan, et al.
Published: (2025)
by: He, Chuan, et al.
Published: (2025)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
by: Kravatskiy, Alexey, et al.
Published: (2025)
by: Kravatskiy, Alexey, et al.
Published: (2025)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
by: Park, Jin Hyun
Published: (2022)
by: Park, Jin Hyun
Published: (2022)
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
by: Sander, Jacob, et al.
Published: (2026)
by: Sander, Jacob, et al.
Published: (2026)
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
MuonRec: Shifting the Optimizer Paradigm Beyond Adam in Scalable Generative Recommendation
by: Shan, Rong, et al.
Published: (2026)
by: Shan, Rong, et al.
Published: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Bootstrap SGD: Algorithmic Stability and Robustness
by: Christmann, Andreas, et al.
Published: (2024)
by: Christmann, Andreas, et al.
Published: (2024)
Adaptive Accountability in Networked MAS: Tracing and Mitigating Emergent Norms at Scale
by: Alqithami, Saad
Published: (2025)
by: Alqithami, Saad
Published: (2025)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
by: Feng, Ce, et al.
Published: (2024)
by: Feng, Ce, et al.
Published: (2024)
Fuzzy Representation of Norms
by: Assadi, Ziba, et al.
Published: (2026)
by: Assadi, Ziba, et al.
Published: (2026)
Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
by: Duhme, Christof, et al.
Published: (2026)
by: Duhme, Christof, et al.
Published: (2026)
Accumulative SGD Influence Estimation for Data Attribution
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization
by: Cui, Chaoqun, et al.
Published: (2026)
by: Cui, Chaoqun, et al.
Published: (2026)
Similar Items
-
Turbo-Muon: Accelerating Orthogonality-Based Optimization with Pre-Conditioning
by: Boissin, Thibaut, et al.
Published: (2025) -
Efficient Robust Conformal Prediction via Lipschitz-Bounded Networks
by: Massena, Thomas, et al.
Published: (2025) -
An Adaptive Orthogonal Convolution Scheme for Efficient and Flexible CNN Architectures
by: Boissin, Thibaut, et al.
Published: (2025) -
Fast and Flexible Robustness Certificates for Semantic Segmentation
by: Massena, Thomas, et al.
Published: (2025) -
DP-SGD Without Clipping: The Lipschitz Neural Network Way
by: Bethune, Louis, et al.
Published: (2023)