Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Ke Liang, Marshall, Noah, Agarwala, Atish, Paquette, Elliot |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
On the Byzantine Fault Tolerance of signSGD with Majority Vote
by: Mengoli, Emanuele, et al.
Published: (2025)
by: Mengoli, Emanuele, et al.
Published: (2025)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026)
by: Agarwala, Atish
Published: (2026)
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
by: Collins-Woodfin, Elizabeth, et al.
Published: (2024)
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
by: Everett, Katie, et al.
Published: (2026)
by: Everett, Katie, et al.
Published: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
by: Agarwala, Atish, et al.
Published: (2024)
by: Agarwala, Atish, et al.
Published: (2024)
Power-Law Spectrum of the Random Feature Model
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Neglected Hessian component explains mysteries in Sharpness regularization
by: Dauphin, Yann N., et al.
Published: (2024)
by: Dauphin, Yann N., et al.
Published: (2024)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
by: Roulet, Vincent, et al.
Published: (2023)
by: Roulet, Vincent, et al.
Published: (2023)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
What do near-optimal learning rate schedules look like?
by: Naganuma, Hiroki, et al.
Published: (2026)
by: Naganuma, Hiroki, et al.
Published: (2026)
Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
by: Benigni, Lucas, et al.
Published: (2025)
by: Benigni, Lucas, et al.
Published: (2025)
4+3 Phases of Compute-Optimal Neural Scaling Laws
by: Paquette, Elliot, et al.
Published: (2024)
by: Paquette, Elliot, et al.
Published: (2024)
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
by: Vary, Simon, et al.
Published: (2026)
by: Vary, Simon, et al.
Published: (2026)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Improving Implicit Regularization of SGD with Preconditioning for Least Square Problems
by: Su, Junwei, et al.
Published: (2024)
by: Su, Junwei, et al.
Published: (2024)
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
High-Dimensional Privacy-Utility Dynamics of Noisy Stochastic Gradient Descent on Least Squares
by: Lin, Shurong, et al.
Published: (2025)
by: Lin, Shurong, et al.
Published: (2025)
Logarithmic-time Schedules for Scaling Language Models with Momentum
by: Ferbach, Damien, et al.
Published: (2026)
by: Ferbach, Damien, et al.
Published: (2026)
Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization
by: Liu, Andy Zeyi, et al.
Published: (2026)
by: Liu, Andy Zeyi, et al.
Published: (2026)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Population Risk Bounds for Kolmogorov-Arnold Networks Trained by DP-SGD with Correlated Noise
by: Wang, Puyu, et al.
Published: (2026)
by: Wang, Puyu, et al.
Published: (2026)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
by: Labarrière, Hippolyte, et al.
Published: (2026)
by: Labarrière, Hippolyte, et al.
Published: (2026)
Exact Mean Square Linear Stability Analysis for SGD
by: Mulayoff, Rotem, et al.
Published: (2023)
by: Mulayoff, Rotem, et al.
Published: (2023)
Correlating Cross-Iteration Noise for DP-SGD using Model Curvature
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
How far away are truly hyperparameter-free learning algorithms?
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
Stochastic Dimension Implicit Functional Projections for Exact Integral Conservation in High-Dimensional PINNs
by: Liang, Zhangyong
Published: (2026)
by: Liang, Zhangyong
Published: (2026)
Anisotropic local law for non-separable sample covariance matrices
by: Fan, Zhou, et al.
Published: (2026)
by: Fan, Zhou, et al.
Published: (2026)
Can SGD Handle Heavy-Tailed Noise?
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
Lap2: Revisiting Laplace DP-SGD for High Dimensions via Majorization Theory
by: Mohammady, Meisam, et al.
Published: (2026)
by: Mohammady, Meisam, et al.
Published: (2026)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
by: Vyas, Nikhil, et al.
Published: (2023)
by: Vyas, Nikhil, et al.
Published: (2023)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
by: Xu, Yizhou, et al.
Published: (2026)
by: Xu, Yizhou, et al.
Published: (2026)
On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature
by: Zhang, Yikuan, et al.
Published: (2026)
by: Zhang, Yikuan, et al.
Published: (2026)
Similar Items
-
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024) -
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026) -
On the Byzantine Fault Tolerance of signSGD with Majority Vote
by: Mengoli, Emanuele, et al.
Published: (2025) -
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025) -
High dimensional theory of two-phase optimizers
by: Agarwala, Atish
Published: (2026)