Dimension-adapted Momentum Outscales SGD

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferbach, Damien, Everett, Katie, Gidel, Gauthier, Paquette, Elliot, Paquette, Courtney
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912386599878656
author Ferbach, Damien
Everett, Katie
Gidel, Gauthier
Paquette, Elliot
Paquette, Courtney
author_facet Ferbach, Damien
Everett, Katie
Gidel, Gauthier
Paquette, Elliot
Paquette, Courtney
contents We investigate scaling laws for stochastic momentum algorithms with small batch on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying data-target complexities. While traditional stochastic gradient descent with momentum (SGD-M) yields identical scaling law exponents to SGD, dimension-adapted Nesterov acceleration (DANA) improves these exponents by scaling momentum hyperparameters based on model size and data complexity. This outscaling phenomenon, which also improves compute-optimal scaling behavior, is achieved by DANA across a broad range of data and target complexities, while traditional methods fall short. Extensive experiments on high-dimensional synthetic quadratics validate our theoretical predictions and large-scale text experiments with LSTMs show DANA's improved loss exponents over SGD hold in a practical setting.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16098
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dimension-adapted Momentum Outscales SGD
Ferbach, Damien
Everett, Katie
Gidel, Gauthier
Paquette, Elliot
Paquette, Courtney
Machine Learning
Optimization and Control
We investigate scaling laws for stochastic momentum algorithms with small batch on the power law random features model, parameterized by data complexity, target complexity, and model size. When trained with a stochastic momentum algorithm, our analysis reveals four distinct loss curve shapes determined by varying data-target complexities. While traditional stochastic gradient descent with momentum (SGD-M) yields identical scaling law exponents to SGD, dimension-adapted Nesterov acceleration (DANA) improves these exponents by scaling momentum hyperparameters based on model size and data complexity. This outscaling phenomenon, which also improves compute-optimal scaling behavior, is achieved by DANA across a broad range of data and target complexities, while traditional methods fall short. Extensive experiments on high-dimensional synthetic quadratics validate our theoretical predictions and large-scale text experiments with LSTMs show DANA's improved loss exponents over SGD hold in a practical setting.
title Dimension-adapted Momentum Outscales SGD
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2505.16098