Why is parameter averaging beneficial in SGD? An objective smoothing perspective

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nitanda, Atsushi, Kikuchi, Ryuhei, Maeda, Shugo, Wu, Denny
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916260040671232
author Nitanda, Atsushi
Kikuchi, Ryuhei
Maeda, Shugo
Wu, Denny
author_facet Nitanda, Atsushi
Kikuchi, Ryuhei
Maeda, Shugo
Wu, Denny
contents It is often observed that stochastic gradient descent (SGD) and its variants implicitly select a solution with good generalization performance; such implicit bias is often characterized in terms of the sharpness of the minima. Kleinberg et al. (2018) connected this bias with the smoothing effect of SGD which eliminates sharp local minima by the convolution using the stochastic gradient noise. We follow this line of research and study the commonly-used averaged SGD algorithm, which has been empirically observed in Izmailov et al. (2018) to prefer a flat minimum and therefore achieves better generalization. We prove that in certain problem settings, averaged SGD can efficiently optimize the smoothed objective which avoids sharp local minima. In experiments, we verify our theory and show that parameter averaging with an appropriate step size indeed leads to significant improvement in the performance of SGD.
format Preprint
id arxiv_https___arxiv_org_abs_2302_09376
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Why is parameter averaging beneficial in SGD? An objective smoothing perspective
Nitanda, Atsushi
Kikuchi, Ryuhei
Maeda, Shugo
Wu, Denny
Machine Learning
It is often observed that stochastic gradient descent (SGD) and its variants implicitly select a solution with good generalization performance; such implicit bias is often characterized in terms of the sharpness of the minima. Kleinberg et al. (2018) connected this bias with the smoothing effect of SGD which eliminates sharp local minima by the convolution using the stochastic gradient noise. We follow this line of research and study the commonly-used averaged SGD algorithm, which has been empirically observed in Izmailov et al. (2018) to prefer a flat minimum and therefore achieves better generalization. We prove that in certain problem settings, averaged SGD can efficiently optimize the smoothed objective which avoids sharp local minima. In experiments, we verify our theory and show that parameter averaging with an appropriate step size indeed leads to significant improvement in the performance of SGD.
title Why is parameter averaging beneficial in SGD? An objective smoothing perspective
topic Machine Learning
url https://arxiv.org/abs/2302.09376