The Effect of Mini-Batch Noise on the Implicit Bias of Adam
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cattaneo, Matias D., Shigida, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Implicit Bias of Adam
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
A Rod Flow Model for Adam at the Edge of Stability
von: Regis, Eric, et al.
Veröffentlicht: (2026)
von: Regis, Eric, et al.
Veröffentlicht: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2026)
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
von: Wang, Shuche, et al.
Veröffentlicht: (2025)
von: Wang, Shuche, et al.
Veröffentlicht: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
AdamZ: An Enhanced Optimisation Method for Neural Network Training
von: Zaznov, Ilia, et al.
Veröffentlicht: (2024)
von: Zaznov, Ilia, et al.
Veröffentlicht: (2024)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Self-Certifying Primal-Dual Optimization Proxies for Large-Scale Batch Economic Dispatch
von: Klamkin, Michael, et al.
Veröffentlicht: (2025)
von: Klamkin, Michael, et al.
Veröffentlicht: (2025)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
von: Huang, Yu, et al.
Veröffentlicht: (2026)
von: Huang, Yu, et al.
Veröffentlicht: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
von: Alvo, Matias, et al.
Veröffentlicht: (2026)
von: Alvo, Matias, et al.
Veröffentlicht: (2026)
Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian Noise
von: Blaser, Ethan, et al.
Veröffentlicht: (2024)
von: Blaser, Ethan, et al.
Veröffentlicht: (2024)
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
von: Attia, Amit, et al.
Veröffentlicht: (2024)
von: Attia, Amit, et al.
Veröffentlicht: (2024)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
von: Choquette-Choo, Christopher A., et al.
Veröffentlicht: (2023)
von: Choquette-Choo, Christopher A., et al.
Veröffentlicht: (2023)
Nonlinear Non-Gaussian Density Steering with Input and Noise Channel Mismatch: Sinkhorn with Memory for Solving the Control-affine Schrödinger Bridge Problem
von: Bondar, Georgiy A., et al.
Veröffentlicht: (2026)
von: Bondar, Georgiy A., et al.
Veröffentlicht: (2026)
DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization
von: Leenders, Nick, et al.
Veröffentlicht: (2025)
von: Leenders, Nick, et al.
Veröffentlicht: (2025)
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
von: Farnia, Farzan, et al.
Veröffentlicht: (2026)
von: Farnia, Farzan, et al.
Veröffentlicht: (2026)
Reward Collapse in Aligning Large Language Models
von: Song, Ziang, et al.
Veröffentlicht: (2023)
von: Song, Ziang, et al.
Veröffentlicht: (2023)
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
von: Grontas, Panagiotis D., et al.
Veröffentlicht: (2025)
von: Grontas, Panagiotis D., et al.
Veröffentlicht: (2025)
Implicit Bias of Mirror Flow on Separable Data
von: Pesme, Scott, et al.
Veröffentlicht: (2024)
von: Pesme, Scott, et al.
Veröffentlicht: (2024)
The Algorithm Configuration Problem
von: Iommazzo, Gabriele, et al.
Veröffentlicht: (2024)
von: Iommazzo, Gabriele, et al.
Veröffentlicht: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
von: Sheng, Jiayuan, et al.
Veröffentlicht: (2025)
von: Sheng, Jiayuan, et al.
Veröffentlicht: (2025)
Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
von: Zhao, Hanyang, et al.
Veröffentlicht: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
Stronger Approximation Guarantees for Non-Monotone γ-Weakly DR-Submodular Maximization
von: Jadav, Hareshkumar, et al.
Veröffentlicht: (2026)
von: Jadav, Hareshkumar, et al.
Veröffentlicht: (2026)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
von: Tsiourvas, Asterios, et al.
Veröffentlicht: (2025)
von: Tsiourvas, Asterios, et al.
Veröffentlicht: (2025)
Reinforcement Learning from Human Feedback with Active Queries
von: Ji, Kaixuan, et al.
Veröffentlicht: (2024)
von: Ji, Kaixuan, et al.
Veröffentlicht: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
von: Gautam, Tanmay, et al.
Veröffentlicht: (2024)
von: Gautam, Tanmay, et al.
Veröffentlicht: (2024)
Algorithmic Challenges in Ensuring Fairness at the Time of Decision
von: Salem, Jad, et al.
Veröffentlicht: (2021)
von: Salem, Jad, et al.
Veröffentlicht: (2021)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
Variational Learning is Effective for Large Deep Networks
von: Shen, Yuesong, et al.
Veröffentlicht: (2024)
von: Shen, Yuesong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Implicit Bias of Adam
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023) -
How Memory in Optimization Algorithms Implicitly Modifies the Loss
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025) -
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025) -
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
von: Baek, Beomhan, et al.
Veröffentlicht: (2025) -
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)