Agnostic Learning of General ReLU Activation Using Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929572364156928 |
|---|---|
| author | Awasthi, Pranjal Tang, Alex Vijayaraghavan, Aravindan |
| author_facet | Awasthi, Pranjal Tang, Alex Vijayaraghavan, Aravindan |
| contents | We provide a convergence analysis of gradient descent for the problem of agnostically learning a single ReLU function with moderate bias under Gaussian distributions. Unlike prior work that studies the setting of zero bias, we consider the more challenging scenario when the bias of the ReLU function is non-zero. Our main result establishes that starting from random initialization, in a polynomial number of iterations gradient descent outputs, with high probability, a ReLU function that achieves an error that is within a constant factor of the optimal error of the best ReLU function with moderate bias. We also provide finite sample guarantees, and these techniques generalize to a broader class of marginal distributions beyond Gaussians. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2208_02711 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | Agnostic Learning of General ReLU Activation Using Gradient Descent Awasthi, Pranjal Tang, Alex Vijayaraghavan, Aravindan Machine Learning Data Structures and Algorithms We provide a convergence analysis of gradient descent for the problem of agnostically learning a single ReLU function with moderate bias under Gaussian distributions. Unlike prior work that studies the setting of zero bias, we consider the more challenging scenario when the bias of the ReLU function is non-zero. Our main result establishes that starting from random initialization, in a polynomial number of iterations gradient descent outputs, with high probability, a ReLU function that achieves an error that is within a constant factor of the optimal error of the best ReLU function with moderate bias. We also provide finite sample guarantees, and these techniques generalize to a broader class of marginal distributions beyond Gaussians. |
| title | Agnostic Learning of General ReLU Activation Using Gradient Descent |
| topic | Machine Learning Data Structures and Algorithms |
| url | https://arxiv.org/abs/2208.02711 |