Optimizer-Induced Mode Connectivity: From AdamW to Muon
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Fangzhao, Kim, Sungyoon, Zhang, Erica, Jiang, Yiqi, Pilanci, Mert |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing Neural Network-Based Generative Diffusion Models through Convex Optimization
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
by: Zhang, Erica, et al.
Published: (2024)
by: Zhang, Erica, et al.
Published: (2024)
Convex Relaxations of ReLU Neural Networks Approximate Global Optima in Polynomial Time
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
From Complexity to Clarity: Analytical Expressions of Deep Neural Network Weights via Clifford's Geometric Algebra and Convexity
by: Pilanci, Mert
Published: (2023)
by: Pilanci, Mert
Published: (2023)
Newton Meets Marchenko-Pastur: Massively Parallel Second-Order Optimization with Hessian Sketching and Debiasing
by: Romanov, Elad, et al.
Published: (2024)
by: Romanov, Elad, et al.
Published: (2024)
Optimal Shrinkage for Distributed Second-Order Optimization
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Spectral Adapter: Fine-Tuning in Spectral Space
by: Zhang, Fangzhao, et al.
Published: (2024)
by: Zhang, Fangzhao, et al.
Published: (2024)
Randomized Geometric Algebra Methods for Convex Neural Networks
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Compressing Large Language Models using Low Rank and Low Precision Decomposition
by: Saha, Rajarshi, et al.
Published: (2024)
by: Saha, Rajarshi, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
by: Li, Huan, et al.
Published: (2026)
by: Li, Huan, et al.
Published: (2026)
Adversarial Training of Two-Layer Polynomial and ReLU Activation Networks via Convex Optimization
by: Kuelbs, Daniel, et al.
Published: (2024)
by: Kuelbs, Daniel, et al.
Published: (2024)
On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
by: Li, Huan, et al.
Published: (2025)
by: Li, Huan, et al.
Published: (2025)
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
by: He, Chuan, et al.
Published: (2025)
by: He, Chuan, et al.
Published: (2025)
A Recovery Guarantee for Sparse Neural Networks
by: Fridovich-Keil, Sara, et al.
Published: (2025)
by: Fridovich-Keil, Sara, et al.
Published: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
A Library of Mirrors: Deep Neural Nets in Low Dimensions are Convex Lasso Models with Reflection Features
by: Zeger, Emi, et al.
Published: (2024)
by: Zeger, Emi, et al.
Published: (2024)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
CRONOS: Enhancing Deep Learning with Scalable GPU Accelerated Convex Neural Networks
by: Feng, Miria, et al.
Published: (2024)
by: Feng, Miria, et al.
Published: (2024)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
A Rod Flow Model for Adam at the Edge of Stability
by: Regis, Eric, et al.
Published: (2026)
by: Regis, Eric, et al.
Published: (2026)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
by: Ranganath, Aditya
Published: (2026)
by: Ranganath, Aditya
Published: (2026)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
On the Condition Number Dependency in Bilevel Optimization
by: Chen, Lesi, et al.
Published: (2025)
by: Chen, Lesi, et al.
Published: (2025)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack
by: Gai, Kuo, et al.
Published: (2024)
by: Gai, Kuo, et al.
Published: (2024)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
by: Wasserkrug, Segev, et al.
Published: (2024)
by: Wasserkrug, Segev, et al.
Published: (2024)
Client-Centric Federated Adaptive Optimization
by: Sun, Jianhui, et al.
Published: (2025)
by: Sun, Jianhui, et al.
Published: (2025)
Contextual Distributionally Robust Optimization with Causal and Continuous Structure: An Interpretable and Tractable Approach
by: Zhang, Fenglin, et al.
Published: (2026)
by: Zhang, Fenglin, et al.
Published: (2026)
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
by: Shumaylov, Zakhar, et al.
Published: (2026)
by: Shumaylov, Zakhar, et al.
Published: (2026)
Similar Items
-
Analyzing Neural Network-Based Generative Diffusion Models through Convex Optimization
by: Zhang, Fangzhao, et al.
Published: (2024) -
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
by: Zhang, Erica, et al.
Published: (2024) -
Convex Relaxations of ReLU Neural Networks Approximate Global Optima in Polynomial Time
by: Kim, Sungyoon, et al.
Published: (2024) -
Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
by: Zhang, Fangzhao, et al.
Published: (2024) -
From Complexity to Clarity: Analytical Expressions of Deep Neural Network Weights via Clifford's Geometric Algebra and Convexity
by: Pilanci, Mert
Published: (2023)