Momentum Further Constrains Sharpness at the Edge of Stochastic Stability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Andreyev, Arseniy, Ananthkumar, Advikar, Walden, Marc, Poggio, Tomaso, Beneventano, Pierfrancesco
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908966972293120
author Andreyev, Arseniy
Ananthkumar, Advikar
Walden, Marc
Poggio, Tomaso
Beneventano, Pierfrancesco
author_facet Andreyev, Arseniy
Ananthkumar, Advikar
Walden, Marc
Poggio, Tomaso
Beneventano, Pierfrancesco
contents Recent work suggests that (stochastic) gradient descent self-organizes near an instability boundary, shaping both optimization and the solutions found. Momentum and mini-batch gradients are widely used in practical deep learning optimization, but it remains unclear whether they operate in a comparable regime of instability. We demonstrate that SGD with momentum exhibits an Edge of Stochastic Stability (EoSS)-like regime with batch-size-dependent behavior that cannot be explained by a single momentum-adjusted stability threshold. Batch Sharpness (the expected directional mini-batch curvature) stabilizes in two distinct regimes: at small batch sizes it converges to a lower plateau $2(1-β)/η$, reflecting amplification of stochastic fluctuations by momentum and favoring flatter regions than vanilla SGD; at large batch sizes it converges to a higher plateau $2(1+β)/η$, where momentum recovers its classical stabilizing effect and favors sharper regions consistent with full-batch dynamics. We further show that this aligns with linear stability thresholds and discuss the implications for hyperparameter tuning and coupling.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14108
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
Andreyev, Arseniy
Ananthkumar, Advikar
Walden, Marc
Poggio, Tomaso
Beneventano, Pierfrancesco
Machine Learning
Dynamical Systems
Optimization and Control
Recent work suggests that (stochastic) gradient descent self-organizes near an instability boundary, shaping both optimization and the solutions found. Momentum and mini-batch gradients are widely used in practical deep learning optimization, but it remains unclear whether they operate in a comparable regime of instability. We demonstrate that SGD with momentum exhibits an Edge of Stochastic Stability (EoSS)-like regime with batch-size-dependent behavior that cannot be explained by a single momentum-adjusted stability threshold. Batch Sharpness (the expected directional mini-batch curvature) stabilizes in two distinct regimes: at small batch sizes it converges to a lower plateau $2(1-β)/η$, reflecting amplification of stochastic fluctuations by momentum and favoring flatter regions than vanilla SGD; at large batch sizes it converges to a higher plateau $2(1+β)/η$, where momentum recovers its classical stabilizing effect and favors sharper regions consistent with full-batch dynamics. We further show that this aligns with linear stability thresholds and discuss the implications for hyperparameter tuning and coupling.
title Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
topic Machine Learning
Dynamical Systems
Optimization and Control
url https://arxiv.org/abs/2604.14108