Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates
Fuente:
arXiv
Saved in:
| Main Authors: | Glentis, Athanasios, Li, Dawei, Yau, Chung-Yiu, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026)
by: Yau, Chung-Yiu, et al.
Published: (2026)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Cumulative Learning Rate Adaptation: Revisiting Path-Based Schedules for SGD and Adam
by: Atamna, Asma, et al.
Published: (2025)
by: Atamna, Asma, et al.
Published: (2025)
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Feather: An Elegant Solution to Effective DNN Sparsification
by: Georgoulakis, Athanasios Glentis, et al.
Published: (2023)
by: Georgoulakis, Athanasios Glentis, et al.
Published: (2023)
EMC$^2$: Efficient MCMC Negative Sampling for Contrastive Learning with Global Convergence
by: Yau, Chung-Yiu, et al.
Published: (2024)
by: Yau, Chung-Yiu, et al.
Published: (2024)
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models
by: Wei, Quan, et al.
Published: (2025)
by: Wei, Quan, et al.
Published: (2025)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
Differentially Private SGD Without Clipping Bias: An Error-Feedback Approach
by: Zhang, Xinwei, et al.
Published: (2023)
by: Zhang, Xinwei, et al.
Published: (2023)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
A Stochastic Approximation Approach for Efficient Decentralized Optimization on Random Networks
by: Yau, Chung-Yiu, et al.
Published: (2024)
by: Yau, Chung-Yiu, et al.
Published: (2024)
Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Revisiting LocalSGD and SCAFFOLD: Improved Rates and Missing Analysis
by: Luo, Ruichen, et al.
Published: (2025)
by: Luo, Ruichen, et al.
Published: (2025)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
by: Mukherjee, Sagnik, et al.
Published: (2026)
by: Mukherjee, Sagnik, et al.
Published: (2026)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Private and Fair Machine Learning: Revisiting the Disparate Impact of Differentially Private SGD
by: Demelius, Lea, et al.
Published: (2025)
by: Demelius, Lea, et al.
Published: (2025)
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
by: Naganuma, Hiroki, et al.
Published: (2025)
by: Naganuma, Hiroki, et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
Revisiting Learning Rate Control
by: Henheik, Micha, et al.
Published: (2025)
by: Henheik, Micha, et al.
Published: (2025)
Adam-mini: Use Fewer Learning Rates To Gain More
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
Lap2: Revisiting Laplace DP-SGD for High Dimensions via Majorization Theory
by: Mohammady, Meisam, et al.
Published: (2026)
by: Mohammady, Meisam, et al.
Published: (2026)
Conda: Column-Normalized Adam for Training Large Language Models Faster
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
by: Surjanovic, Nikola, et al.
Published: (2025)
by: Surjanovic, Nikola, et al.
Published: (2025)
Similar Items
-
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026) -
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025) -
Cumulative Learning Rate Adaptation: Revisiting Path-Based Schedules for SGD and Adam
by: Atamna, Asma, et al.
Published: (2025) -
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
by: Glentis, Athanasios, et al.
Published: (2025) -
Feather: An Elegant Solution to Effective DNN Sparsification
by: Georgoulakis, Athanasios Glentis, et al.
Published: (2023)