Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Andreyev, Arseniy, Beneventano, Pierfrancesco
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908735745556480
author Andreyev, Arseniy
Beneventano, Pierfrancesco
author_facet Andreyev, Arseniy
Beneventano, Pierfrancesco
contents Recent findings by Cohen et al., 2021, demonstrate that when training neural networks using full-batch gradient descent with a step size of $η$, the largest eigenvalue $λ_{\max}$ of the full-batch Hessian consistently stabilizes around $2/η$. These results have significant implications for convergence and generalization. This, however, is not the case for mini-batch optimization algorithms, limiting the broader applicabilityof the consequences of these findings. We show mini-batch Stochastic Gradient Descent (SGD) trains in a different regime we term Edge of Stochastic Stability (EoSS). In this regime, what stabilizes at $2/η$ is Batch Sharpness: the expected directional curvature of mini-batch Hessians along their corresponding stochastic gradients. As a consequence $λ_{\max}$ -- which is generally smaller than Batch Sharpness -- is suppressed, aligning with the long-standing empirical observation that smaller batches and larger step sizes favor flatter minima. We further discuss implications for mathematical modeling of SGD trajectories.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
Andreyev, Arseniy
Beneventano, Pierfrancesco
Machine Learning
Optimization and Control
Recent findings by Cohen et al., 2021, demonstrate that when training neural networks using full-batch gradient descent with a step size of $η$, the largest eigenvalue $λ_{\max}$ of the full-batch Hessian consistently stabilizes around $2/η$. These results have significant implications for convergence and generalization. This, however, is not the case for mini-batch optimization algorithms, limiting the broader applicabilityof the consequences of these findings. We show mini-batch Stochastic Gradient Descent (SGD) trains in a different regime we term Edge of Stochastic Stability (EoSS). In this regime, what stabilizes at $2/η$ is Batch Sharpness: the expected directional curvature of mini-batch Hessians along their corresponding stochastic gradients. As a consequence $λ_{\max}$ -- which is generally smaller than Batch Sharpness -- is suppressed, aligning with the long-standing empirical observation that smaller batches and larger step sizes favor flatter minima. We further discuss implications for mathematical modeling of SGD trajectories.
title Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2412.20553