The Implicit Bias of Gradient Descent on Separable Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Soudry, Daniel, Hoffer, Elad, Nacson, Mor Shpigel, Gunasekar, Suriya, Srebro, Nathan
Format: Preprint
Published: 2017
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929560300290048
author Soudry, Daniel
Hoffer, Elad
Nacson, Mor Shpigel
Gunasekar, Suriya
Srebro, Nathan
author_facet Soudry, Daniel
Hoffer, Elad
Nacson, Mor Shpigel
Gunasekar, Suriya
Srebro, Nathan
contents We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets. We show the predictor converges to the direction of the max-margin (hard margin SVM) solution. The result also generalizes to other monotone decreasing loss functions with an infimum at infinity, to multi-class problems, and to training a weight layer in a deep network in a certain restricted setting. Furthermore, we show this convergence is very slow, and only logarithmic in the convergence of the loss itself. This can help explain the benefit of continuing to optimize the logistic or cross-entropy loss even after the training error is zero and the training loss is extremely small, and, as we show, even if the validation loss increases. Our methodology can also aid in understanding implicit regularization n more complex models and with other optimization methods.
format Preprint
id arxiv_https___arxiv_org_abs_1710_10345
institution arXiv
publishDate 2017
record_format arxiv
spellingShingle The Implicit Bias of Gradient Descent on Separable Data
Soudry, Daniel
Hoffer, Elad
Nacson, Mor Shpigel
Gunasekar, Suriya
Srebro, Nathan
Machine Learning
We examine gradient descent on unregularized logistic regression problems, with homogeneous linear predictors on linearly separable datasets. We show the predictor converges to the direction of the max-margin (hard margin SVM) solution. The result also generalizes to other monotone decreasing loss functions with an infimum at infinity, to multi-class problems, and to training a weight layer in a deep network in a certain restricted setting. Furthermore, we show this convergence is very slow, and only logarithmic in the convergence of the loss itself. This can help explain the benefit of continuing to optimize the logistic or cross-entropy loss even after the training error is zero and the training loss is extremely small, and, as we show, even if the validation loss increases. Our methodology can also aid in understanding implicit regularization n more complex models and with other optimization methods.
title The Implicit Bias of Gradient Descent on Separable Data
topic Machine Learning
url https://arxiv.org/abs/1710.10345