Global Convergence of SGD On Two Layer Neural Nets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gopalani, Pulkit, Mukherjee, Anirbit
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909436821372928
author Gopalani, Pulkit
Mukherjee, Anirbit
author_facet Gopalani, Pulkit
Mukherjee, Anirbit
contents In this note, we consider appropriately regularized $\ell_2-$empirical risk of depth $2$ nets with any number of gates and show bounds on how the empirical loss evolves for SGD iterates on it -- for arbitrary data and if the activation is adequately smooth and bounded like sigmoid and tanh. This in turn leads to a proof of global convergence of SGD for a special class of initializations. We also prove an exponentially fast convergence rate for continuous time SGD that also applies to smooth unbounded activations like SoftPlus. Our key idea is to show the existence of Frobenius norm regularized loss functions on constant-sized neural nets which are "Villani functions" and thus be able to build on recent progress with analyzing SGD on such objectives. Most critically the amount of regularization required for our analysis is independent of the size of the net.
format Preprint
id arxiv_https___arxiv_org_abs_2210_11452
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Global Convergence of SGD On Two Layer Neural Nets
Gopalani, Pulkit
Mukherjee, Anirbit
Machine Learning
Optimization and Control
In this note, we consider appropriately regularized $\ell_2-$empirical risk of depth $2$ nets with any number of gates and show bounds on how the empirical loss evolves for SGD iterates on it -- for arbitrary data and if the activation is adequately smooth and bounded like sigmoid and tanh. This in turn leads to a proof of global convergence of SGD for a special class of initializations. We also prove an exponentially fast convergence rate for continuous time SGD that also applies to smooth unbounded activations like SoftPlus. Our key idea is to show the existence of Frobenius norm regularized loss functions on constant-sized neural nets which are "Villani functions" and thus be able to build on recent progress with analyzing SGD on such objectives. Most critically the amount of regularization required for our analysis is independent of the size of the net.
title Global Convergence of SGD On Two Layer Neural Nets
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2210.11452