High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jagannath, Aukosh, Jones-McCormick, Taj, Sarangian, Varnan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917278591746048
author Jagannath, Aukosh
Jones-McCormick, Taj
Sarangian, Varnan
author_facet Jagannath, Aukosh
Jones-McCormick, Taj
Sarangian, Varnan
contents We develop a high-dimensional scaling limit for Stochastic Gradient Descent with Polyak Momentum (SGD-M) and adaptive step-sizes. This provides a framework to rigourously compare online SGD with some of its popular variants. We show that the scaling limits of SGD-M coincide with those of online SGD after an appropriate time rescaling and a specific choice of step-size. However, if the step-size is kept the same between the two algorithms, SGD-M will amplify high-dimensional effects, potentially degrading performance relative to online SGD. We demonstrate our framework on two popular learning problems: Spiked Tensor PCA and Single Index Models. In both cases, we also examine online SGD with an adaptive step-size based on normalized gradients. In the high-dimensional regime, this algorithm yields multiple benefits: its dynamics admit fixed points closer to the population minimum and widens the range of admissible step-sizes for which the iterates converge to such solutions. These examples provide a rigorous account, aligning with empirical motivation, of how early preconditioners can stabilize and improve dynamics in settings where online SGD fails.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes
Jagannath, Aukosh
Jones-McCormick, Taj
Sarangian, Varnan
Machine Learning
We develop a high-dimensional scaling limit for Stochastic Gradient Descent with Polyak Momentum (SGD-M) and adaptive step-sizes. This provides a framework to rigourously compare online SGD with some of its popular variants. We show that the scaling limits of SGD-M coincide with those of online SGD after an appropriate time rescaling and a specific choice of step-size. However, if the step-size is kept the same between the two algorithms, SGD-M will amplify high-dimensional effects, potentially degrading performance relative to online SGD. We demonstrate our framework on two popular learning problems: Spiked Tensor PCA and Single Index Models. In both cases, we also examine online SGD with an adaptive step-size based on normalized gradients. In the high-dimensional regime, this algorithm yields multiple benefits: its dynamics admit fixed points closer to the population minimum and widens the range of admissible step-sizes for which the iterates converge to such solutions. These examples provide a rigorous account, aligning with empirical motivation, of how early preconditioners can stabilize and improve dynamics in settings where online SGD fails.
title High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes
topic Machine Learning
url https://arxiv.org/abs/2511.03952