Statistical Guarantees for Approximate Stationary Points of Shallow Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Taheri, Mahsa, Xie, Fang, Lederer, Johannes
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912750114963456
author Taheri, Mahsa
Xie, Fang
Lederer, Johannes
author_facet Taheri, Mahsa
Xie, Fang
Lederer, Johannes
contents Since statistical guarantees for neural networks are usually restricted to global optima of intricate objective functions, it is unclear whether these theories explain the performances of actual outputs of neural network pipelines. The goal of this paper is, therefore, to bring statistical theory closer to practice. We develop statistical guarantees for shallow linear neural networks that coincide up to logarithmic factors with the global optima but apply to stationary points and the points nearby. These results support the common notion that neural networks do not necessarily need to be optimized globally from a mathematical perspective. We then extend our statistical guarantees to shallow ReLU neural networks, assuming the first layer weight matrices are nearly identical for the stationary network and the target. More generally, despite being limited to shallow neural networks for now, our theories make an important step forward in describing the practical properties of neural networks in mathematical terms.
format Preprint
id arxiv_https___arxiv_org_abs_2205_04491
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Statistical Guarantees for Approximate Stationary Points of Shallow Neural Networks
Taheri, Mahsa
Xie, Fang
Lederer, Johannes
Machine Learning
Statistics Theory
Since statistical guarantees for neural networks are usually restricted to global optima of intricate objective functions, it is unclear whether these theories explain the performances of actual outputs of neural network pipelines. The goal of this paper is, therefore, to bring statistical theory closer to practice. We develop statistical guarantees for shallow linear neural networks that coincide up to logarithmic factors with the global optima but apply to stationary points and the points nearby. These results support the common notion that neural networks do not necessarily need to be optimized globally from a mathematical perspective. We then extend our statistical guarantees to shallow ReLU neural networks, assuming the first layer weight matrices are nearly identical for the stationary network and the target. More generally, despite being limited to shallow neural networks for now, our theories make an important step forward in describing the practical properties of neural networks in mathematical terms.
title Statistical Guarantees for Approximate Stationary Points of Shallow Neural Networks
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2205.04491