The Implicit Bias of Adam on Separable Data
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chenyang, Zou, Difan, Cao, Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
by: Zhang, Chenyang, et al.
Published: (2025)
by: Zhang, Chenyang, et al.
Published: (2025)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
Improving Implicit Regularization of SGD with Preconditioning for Least Square Problems
by: Su, Junwei, et al.
Published: (2024)
by: Su, Junwei, et al.
Published: (2024)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Initialization Matters: On the Benign Overfitting of Two-Layer ReLU CNN with Fully Trainable Layers
by: Shang, Shuning, et al.
Published: (2024)
by: Shang, Shuning, et al.
Published: (2024)
On the Feature Learning in Diffusion Models
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference
by: Han, Yujin, et al.
Published: (2024)
by: Han, Yujin, et al.
Published: (2024)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
by: Zhang, Chenyang, et al.
Published: (2026)
by: Zhang, Chenyang, et al.
Published: (2026)
Adam Simplified: Bias Correction Debunked
by: Laing, Sam, et al.
Published: (2025)
by: Laing, Sam, et al.
Published: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
by: Hu, Yunzhe, et al.
Published: (2024)
by: Hu, Yunzhe, et al.
Published: (2024)
PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks
by: Su, Junwei, et al.
Published: (2024)
by: Su, Junwei, et al.
Published: (2024)
Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory
by: Bai, Hanru, et al.
Published: (2025)
by: Bai, Hanru, et al.
Published: (2025)
Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow
by: Bai, Hanru, et al.
Published: (2026)
by: Bai, Hanru, et al.
Published: (2026)
On the Limitation and Experience Replay for GNNs in Continual Learning
by: Su, Junwei, et al.
Published: (2023)
by: Su, Junwei, et al.
Published: (2023)
Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
by: Hu, Yunzhe, et al.
Published: (2025)
by: Hu, Yunzhe, et al.
Published: (2025)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
by: Zhang, Ji, et al.
Published: (2023)
by: Zhang, Ji, et al.
Published: (2023)
Effective Blind Source Separation Based on the Adam Algorithm
by: Scarpiniti, Michele, et al.
Published: (2016)
by: Scarpiniti, Michele, et al.
Published: (2016)
Towards Robust Graph Incremental Learning on Evolving Graphs
by: Su, Junwei, et al.
Published: (2024)
by: Su, Junwei, et al.
Published: (2024)
On the Benefits of Over-parameterization for Out-of-Distribution Generalization
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
F-Adapter: Frequency-Adaptive Parameter-Efficient Fine-Tuning in Scientific Machine Learning
by: Zhang, Hangwei, et al.
Published: (2025)
by: Zhang, Hangwei, et al.
Published: (2025)
The Implicit Bias of Logit Regularization
by: Beck, Alon, et al.
Published: (2026)
by: Beck, Alon, et al.
Published: (2026)
Similar Items
-
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
by: Zhang, Chenyang, et al.
Published: (2025) -
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
by: Tang, Xuan, et al.
Published: (2025) -
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026) -
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023) -
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)