A Minimalist Example of Edge-of-Stability and Progressive Sharpening

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Liming, Zhang, Zixuan, Du, Simon, Zhao, Tuo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909524935311360
author Liu, Liming
Zhang, Zixuan
Du, Simon
Zhao, Tuo
author_facet Liu, Liming
Zhang, Zixuan
Du, Simon
Zhao, Tuo
contents Recent advances in deep learning optimization have unveiled two intriguing phenomena under large learning rates: Edge of Stability (EoS) and Progressive Sharpening (PS), challenging classical Gradient Descent (GD) analyses. Current research approaches, using either generalist frameworks or minimalist examples, face significant limitations in explaining these phenomena. This paper advances the minimalist approach by introducing a two-layer network with a two-dimensional input, where one dimension is relevant to the response and the other is irrelevant. Through this model, we rigorously prove the existence of progressive sharpening and self-stabilization under large learning rates, and establish non-asymptotic analysis of the training dynamics and sharpness along the entire GD trajectory. Besides, we connect our minimalist example to existing works by reconciling the existence of a well-behaved ``stable set" between minimalist and generalist analyses, and extending the analysis of Gradient Flow Solution sharpness to our two-dimensional input scenario. These findings provide new insights into the EoS phenomenon from both parameter and input data distribution perspectives, potentially informing more effective optimization strategies in deep learning practice.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02809
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Minimalist Example of Edge-of-Stability and Progressive Sharpening
Liu, Liming
Zhang, Zixuan
Du, Simon
Zhao, Tuo
Machine Learning
Recent advances in deep learning optimization have unveiled two intriguing phenomena under large learning rates: Edge of Stability (EoS) and Progressive Sharpening (PS), challenging classical Gradient Descent (GD) analyses. Current research approaches, using either generalist frameworks or minimalist examples, face significant limitations in explaining these phenomena. This paper advances the minimalist approach by introducing a two-layer network with a two-dimensional input, where one dimension is relevant to the response and the other is irrelevant. Through this model, we rigorously prove the existence of progressive sharpening and self-stabilization under large learning rates, and establish non-asymptotic analysis of the training dynamics and sharpness along the entire GD trajectory. Besides, we connect our minimalist example to existing works by reconciling the existence of a well-behaved ``stable set" between minimalist and generalist analyses, and extending the analysis of Gradient Flow Solution sharpness to our two-dimensional input scenario. These findings provide new insights into the EoS phenomenon from both parameter and input data distribution perspectives, potentially informing more effective optimization strategies in deep learning practice.
title A Minimalist Example of Edge-of-Stability and Progressive Sharpening
topic Machine Learning
url https://arxiv.org/abs/2503.02809