Sharpness-Aware Minimization and the Edge of Stability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Philip M., Bartlett, Peter L.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911906926690304
author Long, Philip M.
Bartlett, Peter L.
author_facet Long, Philip M.
Bartlett, Peter L.
contents Recent experiments have shown that, often, when training a neural network with gradient descent (GD) with a step size $η$, the operator norm of the Hessian of the loss grows until it approximately reaches $2/η$, after which it fluctuates around this value. The quantity $2/η$ has been called the "edge of stability" based on consideration of a local quadratic approximation of the loss. We perform a similar calculation to arrive at an "edge of stability" for Sharpness-Aware Minimization (SAM), a variant of GD which has been shown to improve its generalization. Unlike the case for GD, the resulting SAM-edge depends on the norm of the gradient. Using three deep learning training tasks, we see empirically that SAM operates on the edge of stability identified by this analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2309_12488
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sharpness-Aware Minimization and the Edge of Stability
Long, Philip M.
Bartlett, Peter L.
Machine Learning
Neural and Evolutionary Computing
Recent experiments have shown that, often, when training a neural network with gradient descent (GD) with a step size $η$, the operator norm of the Hessian of the loss grows until it approximately reaches $2/η$, after which it fluctuates around this value. The quantity $2/η$ has been called the "edge of stability" based on consideration of a local quadratic approximation of the loss. We perform a similar calculation to arrive at an "edge of stability" for Sharpness-Aware Minimization (SAM), a variant of GD which has been shown to improve its generalization. Unlike the case for GD, the resulting SAM-edge depends on the norm of the gradient. Using three deep learning training tasks, we see empirically that SAM operates on the edge of stability identified by this analysis.
title Sharpness-Aware Minimization and the Edge of Stability
topic Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2309.12488