Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Sahu, Sharan, Sarkar, Abir, Hogan, Cameron J., Wells, Martin T. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
by: Sahu, Sharan, et al.
Published: (2025)
by: Sahu, Sharan, et al.
Published: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Double Local-to-Unity: Inference under Nearly Nonstationary Volatility
by: Sarkar, Abir, et al.
Published: (2025)
by: Sarkar, Abir, et al.
Published: (2025)
Provably Reliable Classifier Guidance via Cross-Entropy Control
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025)
by: Sahu, Sharan
Published: (2025)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Replay Can Provably Increase Forgetting
by: Mahdaviyeh, Yasaman, et al.
Published: (2025)
by: Mahdaviyeh, Yasaman, et al.
Published: (2025)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025)
by: Vaswani, Sharan, et al.
Published: (2025)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
by: Modoranu, Ionut-Vlad, et al.
Published: (2024)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
by: Dronen, Nicholas, et al.
Published: (2025)
by: Dronen, Nicholas, et al.
Published: (2025)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
by: Vaswani, Sharan, et al.
Published: (2026)
by: Vaswani, Sharan, et al.
Published: (2026)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Cumulative Learning Rate Adaptation: Revisiting Path-Based Schedules for SGD and Adam
by: Atamna, Asma, et al.
Published: (2025)
by: Atamna, Asma, et al.
Published: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Provable Adaptivity of Adam under Non-uniform Smoothness
by: Wang, Bohan, et al.
Published: (2022)
by: Wang, Bohan, et al.
Published: (2022)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
by: Ye, Qilin, et al.
Published: (2025)
by: Ye, Qilin, et al.
Published: (2025)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
by: Jiang, Ruichen, et al.
Published: (2024)
by: Jiang, Ruichen, et al.
Published: (2024)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
by: Glentis, Athanasios, et al.
Published: (2026)
by: Glentis, Athanasios, et al.
Published: (2026)
Cooperative SGD with Dynamic Mixing Matrices
by: Sarkar, Soumya, et al.
Published: (2025)
by: Sarkar, Soumya, et al.
Published: (2025)
Generalization and Optimization of SGD with Lookahead
by: Li, Kangcheng, et al.
Published: (2025)
by: Li, Kangcheng, et al.
Published: (2025)
Adaptive Linear Embedding for Nonstationary High-Dimensional Optimization
by: Wen, Yuejiang, et al.
Published: (2025)
by: Wen, Yuejiang, et al.
Published: (2025)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
by: Mukherjee, Sagnik, et al.
Published: (2026)
by: Mukherjee, Sagnik, et al.
Published: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
First Provably Optimal Asynchronous SGD for Homogeneous and Heterogeneous Data
by: Maranjyan, Artavazd
Published: (2026)
by: Maranjyan, Artavazd
Published: (2026)
Is There an AI Bubble? Robust Date-Stamping for Periods of Exuberance
by: Sarkar, Abir, et al.
Published: (2026)
by: Sarkar, Abir, et al.
Published: (2026)
Structured and Fast Optimization: The Kronecker SGD Algorithm
by: Song, Zhao, et al.
Published: (2023)
by: Song, Zhao, et al.
Published: (2023)
Balancing Utility and Privacy: Dynamically Private SGD with Random Projection
by: Jiang, Zhanhong, et al.
Published: (2025)
by: Jiang, Zhanhong, et al.
Published: (2025)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Quick-Draw Bandits: Quickly Optimizing in Nonstationary Environments with Extremely Many Arms
by: Everett, Derek, et al.
Published: (2025)
by: Everett, Derek, et al.
Published: (2025)
The Optimization Landscape of SGD Across the Feature Learning Strength
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
Nonstationary Sparse Spectral Permanental Process
by: Sun, Zicheng, et al.
Published: (2024)
by: Sun, Zicheng, et al.
Published: (2024)
Accuracy vs. Accuracy: Computational Tradeoffs Between Classification Rates and Utility
by: Amit, Noga, et al.
Published: (2025)
by: Amit, Noga, et al.
Published: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
by: Marek, Martin, et al.
Published: (2026)
by: Marek, Martin, et al.
Published: (2026)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Similar Items
-
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026) -
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
by: Sahu, Sharan, et al.
Published: (2025) -
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025) -
Double Local-to-Unity: Inference under Nearly Nonstationary Volatility
by: Sarkar, Abir, et al.
Published: (2025) -
Provably Reliable Classifier Guidance via Cross-Entropy Control
by: Sahu, Sharan, et al.
Published: (2026)