Agnostic Learning of General ReLU Activation Using Gradient Descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Awasthi, Pranjal, Tang, Alex, Vijayaraghavan, Aravindan
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929572364156928
author Awasthi, Pranjal
Tang, Alex
Vijayaraghavan, Aravindan
author_facet Awasthi, Pranjal
Tang, Alex
Vijayaraghavan, Aravindan
contents We provide a convergence analysis of gradient descent for the problem of agnostically learning a single ReLU function with moderate bias under Gaussian distributions. Unlike prior work that studies the setting of zero bias, we consider the more challenging scenario when the bias of the ReLU function is non-zero. Our main result establishes that starting from random initialization, in a polynomial number of iterations gradient descent outputs, with high probability, a ReLU function that achieves an error that is within a constant factor of the optimal error of the best ReLU function with moderate bias. We also provide finite sample guarantees, and these techniques generalize to a broader class of marginal distributions beyond Gaussians.
format Preprint
id arxiv_https___arxiv_org_abs_2208_02711
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Agnostic Learning of General ReLU Activation Using Gradient Descent
Awasthi, Pranjal
Tang, Alex
Vijayaraghavan, Aravindan
Machine Learning
Data Structures and Algorithms
We provide a convergence analysis of gradient descent for the problem of agnostically learning a single ReLU function with moderate bias under Gaussian distributions. Unlike prior work that studies the setting of zero bias, we consider the more challenging scenario when the bias of the ReLU function is non-zero. Our main result establishes that starting from random initialization, in a polynomial number of iterations gradient descent outputs, with high probability, a ReLU function that achieves an error that is within a constant factor of the optimal error of the best ReLU function with moderate bias. We also provide finite sample guarantees, and these techniques generalize to a broader class of marginal distributions beyond Gaussians.
title Agnostic Learning of General ReLU Activation Using Gradient Descent
topic Machine Learning
Data Structures and Algorithms
url https://arxiv.org/abs/2208.02711