The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Gronich, Eitan, Vardi, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022)
by: Timor, Nadav, et al.
Published: (2022)
ImpMIA: Leveraging Implicit Bias for Membership Inference Attack
by: Golbari, Yuval, et al.
Published: (2025)
by: Golbari, Yuval, et al.
Published: (2025)
Provable Privacy Attacks on Trained Shallow Neural Networks
by: Smorodinsky, Guy, et al.
Published: (2024)
by: Smorodinsky, Guy, et al.
Published: (2024)
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)
by: Zhang, Chenyang, et al.
Published: (2024)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Provable Unlearning with Gradient Ascent on Two-Layer ReLU Neural Networks
by: Melamed, Odelia, et al.
Published: (2025)
by: Melamed, Odelia, et al.
Published: (2025)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
by: Joshi, Nirmit, et al.
Published: (2023)
by: Joshi, Nirmit, et al.
Published: (2023)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024)
by: Frei, Spencer, et al.
Published: (2024)
Transformers are almost optimal metalearners for linear classification
by: Magen, Roey, et al.
Published: (2025)
by: Magen, Roey, et al.
Published: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Querying Kernel Methods Suffices for Reconstructing their Training Data
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024)
by: Medvedev, Marko, et al.
Published: (2024)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Can Muon Fine-tune Adam-Pretrained Models?
by: Qu, Xingyu, et al.
Published: (2026)
by: Qu, Xingyu, et al.
Published: (2026)
Implicit Graph Neural Diffusion Networks: Convergence, Generalization, and Over-Smoothing
by: Fu, Guoji, et al.
Published: (2023)
by: Fu, Guoji, et al.
Published: (2023)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
by: Vyas, Nikhil, et al.
Published: (2023)
by: Vyas, Nikhil, et al.
Published: (2023)
Implicit Bias of Mirror Flow for Shallow Neural Networks in Univariate Regression
by: Liang, Shuang, et al.
Published: (2024)
by: Liang, Shuang, et al.
Published: (2024)
Factorized Neural Implicit DMD for Parametric Dynamics
by: Chen, Siyuan, et al.
Published: (2026)
by: Chen, Siyuan, et al.
Published: (2026)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
by: Zhou, Lijia, et al.
Published: (2023)
by: Zhou, Lijia, et al.
Published: (2023)
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Adam Simplified: Bias Correction Debunked
by: Laing, Sam, et al.
Published: (2025)
by: Laing, Sam, et al.
Published: (2025)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
Quotient Geometry, Effective Curvature, and Implicit Bias in Simple Shallow Neural Networks
by: Dong, Hang-Cheng, et al.
Published: (2026)
by: Dong, Hang-Cheng, et al.
Published: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
by: Dragutinović, Sara, et al.
Published: (2026)
by: Dragutinović, Sara, et al.
Published: (2026)
Provable Adaptivity of Adam under Non-uniform Smoothness
by: Wang, Bohan, et al.
Published: (2022)
by: Wang, Bohan, et al.
Published: (2022)
The Implicit Bias of Logit Regularization
by: Beck, Alon, et al.
Published: (2026)
by: Beck, Alon, et al.
Published: (2026)
Positive Distribution Shift as a Framework for Understanding Tractable Learning
by: Medvedev, Marko, et al.
Published: (2026)
by: Medvedev, Marko, et al.
Published: (2026)
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024)
by: Magen, Roey, et al.
Published: (2024)
How Does Overparameterization Affect Machine Unlearning of Deep Neural Networks?
by: Alon, Gal, et al.
Published: (2025)
by: Alon, Gal, et al.
Published: (2025)
Similar Items
-
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024) -
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022) -
ImpMIA: Leveraging Implicit Bias for Membership Inference Attack
by: Golbari, Yuval, et al.
Published: (2025) -
Provable Privacy Attacks on Trained Shallow Neural Networks
by: Smorodinsky, Guy, et al.
Published: (2024) -
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)