High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes
Fuente:
arXiv
Saved in:
| Main Authors: | Jagannath, Aukosh, Jones-McCormick, Taj, Sarangian, Varnan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index Models
by: Jones-McCormick, Taj, et al.
Published: (2025)
by: Jones-McCormick, Taj, et al.
Published: (2025)
Detecting Metastable Basins in High Dimensions via Marginal Trajectory Distribution Discrimination
by: Jones-McCormick, Taj
Published: (2026)
by: Jones-McCormick, Taj
Published: (2026)
Universality of high-dimensional scaling limits of stochastic gradient descent
by: Gheissari, Reza, et al.
Published: (2025)
by: Gheissari, Reza, et al.
Published: (2025)
Optimality of Message-Passing Architectures for Sparse Graphs
by: Baranwal, Aseem, et al.
Published: (2023)
by: Baranwal, Aseem, et al.
Published: (2023)
Spectral alignment of stochastic gradient descent for high-dimensional classification tasks
by: Arous, Gerard Ben, et al.
Published: (2023)
by: Arous, Gerard Ben, et al.
Published: (2023)
Local geometry of high-dimensional mixture models: Effective spectral theory and dynamical transitions
by: Arous, Gerard Ben, et al.
Published: (2025)
by: Arous, Gerard Ben, et al.
Published: (2025)
Differentially private multivariate medians
by: Ramsay, Kelly, et al.
Published: (2022)
by: Ramsay, Kelly, et al.
Published: (2022)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
Adaptive Active Learning for Regression via Reinforcement Learning
by: Nguyen, Simon D., et al.
Published: (2026)
by: Nguyen, Simon D., et al.
Published: (2026)
Unique Rashomon Sets for Robust Active Learning
by: Nguyen, Simon, et al.
Published: (2025)
by: Nguyen, Simon, et al.
Published: (2025)
Dimension-adapted Momentum Outscales SGD
by: Ferbach, Damien, et al.
Published: (2025)
by: Ferbach, Damien, et al.
Published: (2025)
Ordered Momentum for Asynchronous SGD
by: Shi, Chang-Wei, et al.
Published: (2024)
by: Shi, Chang-Wei, et al.
Published: (2024)
Spatially Robust Inference with Predicted and Missing at Random Labels
by: Salerno, Stephen, et al.
Published: (2026)
by: Salerno, Stephen, et al.
Published: (2026)
Adaptive debiased SGD in high-dimensional GLMs with streaming data
by: Han, Ruijian, et al.
Published: (2024)
by: Han, Ruijian, et al.
Published: (2024)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
by: Garg, Sachin, et al.
Published: (2026)
by: Garg, Sachin, et al.
Published: (2026)
A Unified Framework for Inference with General Missingness Patterns and Machine Learning Imputation
by: Chen, Xingran, et al.
Published: (2025)
by: Chen, Xingran, et al.
Published: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
by: Wang, Runzhe, et al.
Published: (2023)
by: Wang, Runzhe, et al.
Published: (2023)
Signal Processing Meets SGD: From Momentum to Filter
by: Yao, Zhipeng, et al.
Published: (2023)
by: Yao, Zhipeng, et al.
Published: (2023)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
by: Sahu, Sharan, et al.
Published: (2026)
by: Sahu, Sharan, et al.
Published: (2026)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
by: Lee, Jason D., et al.
Published: (2024)
by: Lee, Jason D., et al.
Published: (2024)
Spectral goodness-of-fit tests for complete and partial network data
by: Lubold, Shane, et al.
Published: (2021)
by: Lubold, Shane, et al.
Published: (2021)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
Perturbed Iterate SGD for Lipschitz Continuous Loss Functions with Numerical Error and Adaptive Step Sizes
by: Metel, Michael R.
Published: (2022)
by: Metel, Michael R.
Published: (2022)
Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance
by: Oikonomou, Dimitris, et al.
Published: (2024)
by: Oikonomou, Dimitris, et al.
Published: (2024)
Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum
by: Dahan, Tehila, et al.
Published: (2026)
by: Dahan, Tehila, et al.
Published: (2026)
Pseudo-Maximum Likelihood Theory for High-Dimensional Rank One Inference
by: Grant, Curtis, et al.
Published: (2025)
by: Grant, Curtis, et al.
Published: (2025)
Distributed Learning and Inference Systems: A Networking Perspective
by: Moussa, Hesham G., et al.
Published: (2025)
by: Moussa, Hesham G., et al.
Published: (2025)
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
REALITrees: Rashomon Ensemble Active Learning for Interpretable Trees
by: Nguyen, Simon D., et al.
Published: (2026)
by: Nguyen, Simon D., et al.
Published: (2026)
Time-Series Classification in Smart Manufacturing Systems: An Experimental Evaluation of State-of-the-Art Machine Learning Algorithms
by: Farahani, Mojtaba A., et al.
Published: (2023)
by: Farahani, Mojtaba A., et al.
Published: (2023)
Do We Really Even Need Data?
by: Hoffman, Kentaro, et al.
Published: (2024)
by: Hoffman, Kentaro, et al.
Published: (2024)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Robustly estimating heterogeneity in factorial data using Rashomon Partitions
by: Venkateswaran, Aparajithan, et al.
Published: (2024)
by: Venkateswaran, Aparajithan, et al.
Published: (2024)
Physics-Informed Neural Networks for Electrical Circuit Analysis: Applications in Dielectric Material Modeling
by: Taj, Reyhaneh
Published: (2024)
by: Taj, Reyhaneh
Published: (2024)
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data
by: Salerno, Stephen, et al.
Published: (2025)
by: Salerno, Stephen, et al.
Published: (2025)
Similar Items
-
Provable Benefits of Unsupervised Pre-training and Transfer Learning via Single-Index Models
by: Jones-McCormick, Taj, et al.
Published: (2025) -
Detecting Metastable Basins in High Dimensions via Marginal Trajectory Distribution Discrimination
by: Jones-McCormick, Taj
Published: (2026) -
Universality of high-dimensional scaling limits of stochastic gradient descent
by: Gheissari, Reza, et al.
Published: (2025) -
Optimality of Message-Passing Architectures for Sparse Graphs
by: Baranwal, Aseem, et al.
Published: (2023) -
Spectral alignment of stochastic gradient descent for high-dimensional classification tasks
by: Arous, Gerard Ben, et al.
Published: (2023)