Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Phunyaphibarn, Prin, Lee, Junghyun, Wang, Bohan, Zhang, Huishuai, Yun, Chulhee
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911891506331648
author Phunyaphibarn, Prin
Lee, Junghyun
Wang, Bohan
Zhang, Huishuai
Yun, Chulhee
author_facet Phunyaphibarn, Prin
Lee, Junghyun
Wang, Bohan
Zhang, Huishuai
Yun, Chulhee
contents Although gradient descent with Polyak's momentum is widely used in modern machine and deep learning, a concrete understanding of its effects on the training trajectory remains elusive. In this work, we empirically show that for linear diagonal networks and nonlinear neural networks, momentum gradient descent with a large learning rate displays large catapults, driving the iterates towards much flatter minima than those found by gradient descent. We hypothesize that the large catapult is caused by momentum "prolonging" the self-stabilization effect (Damian et al., 2023). We provide theoretical and empirical support for our hypothesis in a simple toy example and empirical evidence supporting our hypothesis for linear diagonal networks.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15051
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
Phunyaphibarn, Prin
Lee, Junghyun
Wang, Bohan
Zhang, Huishuai
Yun, Chulhee
Machine Learning
Optimization and Control
Although gradient descent with Polyak's momentum is widely used in modern machine and deep learning, a concrete understanding of its effects on the training trajectory remains elusive. In this work, we empirically show that for linear diagonal networks and nonlinear neural networks, momentum gradient descent with a large learning rate displays large catapults, driving the iterates towards much flatter minima than those found by gradient descent. We hypothesize that the large catapult is caused by momentum "prolonging" the self-stabilization effect (Damian et al., 2023). We provide theoretical and empirical support for our hypothesis in a simple toy example and empirical evidence supporting our hypothesis for linear diagonal networks.
title Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2311.15051