Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Sahu, Sharan, Sarkar, Abir, Hogan, Cameron J., Wells, Martin T. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
por: Sahu, Sharan, et al.
Publicado: (2026)
por: Sahu, Sharan, et al.
Publicado: (2026)
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
por: Sahu, Sharan, et al.
Publicado: (2025)
por: Sahu, Sharan, et al.
Publicado: (2025)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
por: Vasudeva, Bhavya, et al.
Publicado: (2025)
por: Vasudeva, Bhavya, et al.
Publicado: (2025)
Double Local-to-Unity: Inference under Nearly Nonstationary Volatility
por: Sarkar, Abir, et al.
Publicado: (2025)
por: Sarkar, Abir, et al.
Publicado: (2025)
Provably Reliable Classifier Guidance via Cross-Entropy Control
por: Sahu, Sharan, et al.
Publicado: (2026)
por: Sahu, Sharan, et al.
Publicado: (2026)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
por: Sahu, Sharan
Publicado: (2025)
por: Sahu, Sharan
Publicado: (2025)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
por: Zhang, Yiheng, et al.
Publicado: (2026)
por: Zhang, Yiheng, et al.
Publicado: (2026)
Replay Can Provably Increase Forgetting
por: Mahdaviyeh, Yasaman, et al.
Publicado: (2025)
por: Mahdaviyeh, Yasaman, et al.
Publicado: (2025)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
por: Vaswani, Sharan, et al.
Publicado: (2025)
por: Vaswani, Sharan, et al.
Publicado: (2025)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
por: Modoranu, Ionut-Vlad, et al.
Publicado: (2024)
por: Modoranu, Ionut-Vlad, et al.
Publicado: (2024)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
por: Dronen, Nicholas, et al.
Publicado: (2025)
por: Dronen, Nicholas, et al.
Publicado: (2025)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
por: Vaswani, Sharan, et al.
Publicado: (2026)
por: Vaswani, Sharan, et al.
Publicado: (2026)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
por: Huang, Feihu, et al.
Publicado: (2026)
por: Huang, Feihu, et al.
Publicado: (2026)
Cumulative Learning Rate Adaptation: Revisiting Path-Based Schedules for SGD and Adam
por: Atamna, Asma, et al.
Publicado: (2025)
por: Atamna, Asma, et al.
Publicado: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
por: Zhu, Hanqing, et al.
Publicado: (2024)
por: Zhu, Hanqing, et al.
Publicado: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
por: Jin, Ruinan, et al.
Publicado: (2024)
por: Jin, Ruinan, et al.
Publicado: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
por: Liu, Bingbin, et al.
Publicado: (2025)
por: Liu, Bingbin, et al.
Publicado: (2025)
Provable Adaptivity of Adam under Non-uniform Smoothness
por: Wang, Bohan, et al.
Publicado: (2022)
por: Wang, Bohan, et al.
Publicado: (2022)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
por: Ye, Qilin, et al.
Publicado: (2025)
por: Ye, Qilin, et al.
Publicado: (2025)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
por: Peng, Hanyang, et al.
Publicado: (2025)
por: Peng, Hanyang, et al.
Publicado: (2025)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
por: Srećković, Teodora, et al.
Publicado: (2025)
por: Srećković, Teodora, et al.
Publicado: (2025)
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization
por: Jiang, Ruichen, et al.
Publicado: (2024)
por: Jiang, Ruichen, et al.
Publicado: (2024)
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
por: Glentis, Athanasios, et al.
Publicado: (2026)
por: Glentis, Athanasios, et al.
Publicado: (2026)
Cooperative SGD with Dynamic Mixing Matrices
por: Sarkar, Soumya, et al.
Publicado: (2025)
por: Sarkar, Soumya, et al.
Publicado: (2025)
Generalization and Optimization of SGD with Lookahead
por: Li, Kangcheng, et al.
Publicado: (2025)
por: Li, Kangcheng, et al.
Publicado: (2025)
Adaptive Linear Embedding for Nonstationary High-Dimensional Optimization
por: Wen, Yuejiang, et al.
Publicado: (2025)
por: Wen, Yuejiang, et al.
Publicado: (2025)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
por: Mukherjee, Sagnik, et al.
Publicado: (2026)
por: Mukherjee, Sagnik, et al.
Publicado: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
por: Jin, Ruinan, et al.
Publicado: (2026)
por: Jin, Ruinan, et al.
Publicado: (2026)
First Provably Optimal Asynchronous SGD for Homogeneous and Heterogeneous Data
por: Maranjyan, Artavazd
Publicado: (2026)
por: Maranjyan, Artavazd
Publicado: (2026)
Is There an AI Bubble? Robust Date-Stamping for Periods of Exuberance
por: Sarkar, Abir, et al.
Publicado: (2026)
por: Sarkar, Abir, et al.
Publicado: (2026)
Structured and Fast Optimization: The Kronecker SGD Algorithm
por: Song, Zhao, et al.
Publicado: (2023)
por: Song, Zhao, et al.
Publicado: (2023)
Balancing Utility and Privacy: Dynamically Private SGD with Random Projection
por: Jiang, Zhanhong, et al.
Publicado: (2025)
por: Jiang, Zhanhong, et al.
Publicado: (2025)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
por: Dahan, Tehila, et al.
Publicado: (2023)
por: Dahan, Tehila, et al.
Publicado: (2023)
Quick-Draw Bandits: Quickly Optimizing in Nonstationary Environments with Extremely Many Arms
por: Everett, Derek, et al.
Publicado: (2025)
por: Everett, Derek, et al.
Publicado: (2025)
The Optimization Landscape of SGD Across the Feature Learning Strength
por: Atanasov, Alexander, et al.
Publicado: (2024)
por: Atanasov, Alexander, et al.
Publicado: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
por: Anisimov, Maksim, et al.
Publicado: (2026)
por: Anisimov, Maksim, et al.
Publicado: (2026)
Nonstationary Sparse Spectral Permanental Process
por: Sun, Zicheng, et al.
Publicado: (2024)
por: Sun, Zicheng, et al.
Publicado: (2024)
Accuracy vs. Accuracy: Computational Tradeoffs Between Classification Rates and Utility
por: Amit, Noga, et al.
Publicado: (2025)
por: Amit, Noga, et al.
Publicado: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
por: Marek, Martin, et al.
Publicado: (2026)
por: Marek, Martin, et al.
Publicado: (2026)
Sign-SGD via Parameter-Free Optimization
por: Medyakov, Daniil, et al.
Publicado: (2025)
por: Medyakov, Daniil, et al.
Publicado: (2025)
Ejemplares similares
-
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
por: Sahu, Sharan, et al.
Publicado: (2026) -
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
por: Sahu, Sharan, et al.
Publicado: (2025) -
The Rich and the Simple: On the Implicit Bias of Adam and SGD
por: Vasudeva, Bhavya, et al.
Publicado: (2025) -
Double Local-to-Unity: Inference under Nearly Nonstationary Volatility
por: Sarkar, Abir, et al.
Publicado: (2025) -
Provably Reliable Classifier Guidance via Cross-Entropy Control
por: Sahu, Sharan, et al.
Publicado: (2026)