Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
Fuente:
arXiv
Saved in:
| Main Authors: | Baek, Beomhan, Song, Minhak, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
A Rod Flow Model for Adam at the Edge of Stability
by: Regis, Eric, et al.
Published: (2026)
by: Regis, Eric, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
AdLoCo: adaptive batching significantly improves communications efficiency and convergence for Large Language Models
by: Kutuzov, Nikolay, et al.
Published: (2025)
by: Kutuzov, Nikolay, et al.
Published: (2025)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Incremental Gradient Descent with Small Epoch Counts is Surprisingly Slow on Ill-Conditioned Problems
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable Improvements
by: Chae, Jiseok, et al.
Published: (2024)
by: Chae, Jiseok, et al.
Published: (2024)
Fundamental Benefit of Alternating Updates in Minimax Optimization
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime
by: Geng, Haoyu, et al.
Published: (2023)
by: Geng, Haoyu, et al.
Published: (2023)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
by: Moon, Chaewon, et al.
Published: (2026)
by: Moon, Chaewon, et al.
Published: (2026)
AdamZ: An Enhanced Optimisation Method for Neural Network Training
by: Zaznov, Ilia, et al.
Published: (2024)
by: Zaznov, Ilia, et al.
Published: (2024)
Boosting K-means for Big Data by Fusing Data Streaming with Global Optimization
by: Mussabayev, Ravil, et al.
Published: (2024)
by: Mussabayev, Ravil, et al.
Published: (2024)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
Federated Distributionally Robust Optimization with Non-Convex Objectives: Algorithm and Analysis
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Data-Driven Portfolio Management for Motion Pictures Industry: A New Data-Driven Optimization Methodology Using a Large Language Model as the Expert
by: Alipour-Vaezi, Mohammad, et al.
Published: (2024)
by: Alipour-Vaezi, Mohammad, et al.
Published: (2024)
A Median Perspective on Unlabeled Data for Out-of-Distribution Detection
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
Frequency-aware Surrogate Modeling With SMT Kernels For Advanced Data Forecasting
by: Gonel, Nicolas, et al.
Published: (2025)
by: Gonel, Nicolas, et al.
Published: (2025)
Data-driven Projection Generation for Efficiently Solving Heterogeneous Quadratic Programming Problems
by: Iwata, Tomoharu, et al.
Published: (2025)
by: Iwata, Tomoharu, et al.
Published: (2025)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Kernel-Free Universum Quadratic Surface Twin Support Vector Machines for Imbalanced Data
by: Moosaei, Hossein, et al.
Published: (2024)
by: Moosaei, Hossein, et al.
Published: (2024)
Geometric Neural Operators (GNPs) for Data-Driven Deep Learning of Non-Euclidean Operators
by: Quackenbush, Blaine, et al.
Published: (2024)
by: Quackenbush, Blaine, et al.
Published: (2024)
Similar Items
-
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025) -
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023) -
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024) -
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023) -
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)