Quantitative CLTs in Deep Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Favaro, Stefano, Hanin, Boris, Marinucci, Domenico, Nourdin, Ivan, Peccati, Giovanni
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929387562074112
author Favaro, Stefano
Hanin, Boris
Marinucci, Domenico
Nourdin, Ivan
Peccati, Giovanni
author_facet Favaro, Stefano
Hanin, Boris
Marinucci, Domenico
Nourdin, Ivan
Peccati, Giovanni
contents We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant $n$. Under mild assumptions on the non-linearity, we obtain quantitative bounds on normal approximations valid at large but finite $n$ and any fixed network depth. Our theorems show both for the finite-dimensional distributions and the entire process, that the distance between a random fully connected network (and its derivatives) to the corresponding infinite width Gaussian process scales like $n^{-γ}$ for $γ>0$, with the exponent depending on the metric used to measure discrepancy. Our bounds are strictly stronger in terms of their dependence on network width than any previously available in the literature; in the one-dimensional case, we also prove that they are optimal, i.e., we establish matching lower bounds.
format Preprint
id arxiv_https___arxiv_org_abs_2307_06092
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Quantitative CLTs in Deep Neural Networks
Favaro, Stefano
Hanin, Boris
Marinucci, Domenico
Nourdin, Ivan
Peccati, Giovanni
Machine Learning
Artificial Intelligence
Probability
We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant $n$. Under mild assumptions on the non-linearity, we obtain quantitative bounds on normal approximations valid at large but finite $n$ and any fixed network depth. Our theorems show both for the finite-dimensional distributions and the entire process, that the distance between a random fully connected network (and its derivatives) to the corresponding infinite width Gaussian process scales like $n^{-γ}$ for $γ>0$, with the exponent depending on the metric used to measure discrepancy. Our bounds are strictly stronger in terms of their dependence on network width than any previously available in the literature; in the one-dimensional case, we also prove that they are optimal, i.e., we establish matching lower bounds.
title Quantitative CLTs in Deep Neural Networks
topic Machine Learning
Artificial Intelligence
Probability
url https://arxiv.org/abs/2307.06092