Dimension-adapted Momentum Outscales SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Ferbach, Damien, Everett, Katie, Gidel, Gauthier, Paquette, Elliot, Paquette, Courtney |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Logarithmic-time Schedules for Scaling Language Models with Momentum
by: Ferbach, Damien, et al.
Published: (2026)
by: Ferbach, Damien, et al.
Published: (2026)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026)
by: Everett, Katie, et al.
Published: (2026)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024)
by: Paquette, Elliot, et al.
Published: (2024)
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
Mirror Descent Algorithms with Nearly Dimension-Independent Rates for Differentially-Private Stochastic Saddle-Point Problems
by: González, Tomás, et al.
Published: (2024)
by: González, Tomás, et al.
Published: (2024)
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
by: Kim, Junhyung Lyle, et al.
Published: (2022)
by: Kim, Junhyung Lyle, et al.
Published: (2022)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
by: Ferbach, Damien, et al.
Published: (2023)
by: Ferbach, Damien, et al.
Published: (2023)
Solving Hidden Monotone Variational Inequalities with Surrogate Losses
by: D'Orazio, Ryan, et al.
Published: (2024)
by: D'Orazio, Ryan, et al.
Published: (2024)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
by: Dang, Anh, et al.
Published: (2024)
by: Dang, Anh, et al.
Published: (2024)
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
by: Gorbunov, Eduard, et al.
Published: (2023)
by: Gorbunov, Eduard, et al.
Published: (2023)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
by: Xiao, Ke Liang, et al.
Published: (2024)
by: Xiao, Ke Liang, et al.
Published: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Making SGD Parameter-Free
by: Carmon, Yair, et al.
Published: (2022)
by: Carmon, Yair, et al.
Published: (2022)
On the Trajectories of SGD Without Replacement
by: Beneventano, Pierfrancesco
Published: (2023)
by: Beneventano, Pierfrancesco
Published: (2023)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
Demystifying SGD with Doubly Stochastic Gradients
by: Kim, Kyurae, et al.
Published: (2024)
by: Kim, Kyurae, et al.
Published: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
SGD with memory: fundamental properties and stochastic acceleration
by: Yarotsky, Dmitry, et al.
Published: (2024)
by: Yarotsky, Dmitry, et al.
Published: (2024)
Does SGD really happen in tiny subspaces?
by: Song, Minhak, et al.
Published: (2024)
by: Song, Minhak, et al.
Published: (2024)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
by: Andreyev, Arseniy, et al.
Published: (2024)
by: Andreyev, Arseniy, et al.
Published: (2024)
From Gradient Clipping to Normalization for Heavy Tailed SGD
by: Hübler, Florian, et al.
Published: (2024)
by: Hübler, Florian, et al.
Published: (2024)
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization
by: Zhang, Haihan, et al.
Published: (2024)
by: Zhang, Haihan, et al.
Published: (2024)
Dual-Delayed Asynchronous SGD for Arbitrarily Heterogeneous Data
by: Wang, Xiaolu, et al.
Published: (2024)
by: Wang, Xiaolu, et al.
Published: (2024)
Faster Convergence of Local SGD for Over-Parameterized Models
by: Qin, Tiancheng, et al.
Published: (2022)
by: Qin, Tiancheng, et al.
Published: (2022)
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024)
by: Sun, Ying, et al.
Published: (2024)
Similar Items
-
Logarithmic-time Schedules for Scaling Language Models with Momentum
by: Ferbach, Damien, et al.
Published: (2026) -
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026) -
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026) -
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024) -
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)