The Marginal Value of Momentum for Small Learning Rate SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Runzhe, Malladi, Sadhika, Wang, Tianhao, Lyu, Kaifeng, Li, Zhiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
A Minibatch-SGD-Based Learning Meta-Policy for Inventory Systems with Myopic Optimal Policy
by: Lyu, Jiameng, et al.
Published: (2024)
by: Lyu, Jiameng, et al.
Published: (2024)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
by: Surjanovic, Nikola, et al.
Published: (2025)
by: Surjanovic, Nikola, et al.
Published: (2025)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025)
by: Kovalev, Dmitry, et al.
Published: (2025)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Sharp High-Probability Rates for Nonlinear SGD under Heavy-Tailed Noise via Symmetrization
by: Armacki, Aleksandar, et al.
Published: (2025)
by: Armacki, Aleksandar, et al.
Published: (2025)
Dual-Delayed Asynchronous SGD for Arbitrarily Heterogeneous Data
by: Wang, Xiaolu, et al.
Published: (2024)
by: Wang, Xiaolu, et al.
Published: (2024)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
Making SGD Parameter-Free
by: Carmon, Yair, et al.
Published: (2022)
by: Carmon, Yair, et al.
Published: (2022)
Byzantine-Robust Distributed SGD: A Unified Analysis and Tight Error Bounds
by: Ruan, Boyuan, et al.
Published: (2026)
by: Ruan, Boyuan, et al.
Published: (2026)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
Demystifying SGD with Doubly Stochastic Gradients
by: Kim, Kyurae, et al.
Published: (2024)
by: Kim, Kyurae, et al.
Published: (2024)
Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
by: Sarkar, Krisanu
Published: (2025)
by: Sarkar, Krisanu
Published: (2025)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
SGD with memory: fundamental properties and stochastic acceleration
by: Yarotsky, Dmitry, et al.
Published: (2024)
by: Yarotsky, Dmitry, et al.
Published: (2024)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
by: Li, Chris Junchi
Published: (2024)
by: Li, Chris Junchi
Published: (2024)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Large Deviation Upper Bounds and Improved MSE Rates of Nonlinear SGD: Heavy-tailed Noise and Power of Symmetry
by: Armacki, Aleksandar, et al.
Published: (2024)
by: Armacki, Aleksandar, et al.
Published: (2024)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
by: Zhang, Haihan, et al.
Published: (2024)
by: Zhang, Haihan, et al.
Published: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
by: Qin, Tiancheng, et al.
Published: (2022)
by: Qin, Tiancheng, et al.
Published: (2022)
Similar Items
-
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025) -
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025) -
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025) -
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026) -
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)