Dimension-independent rates for structured neural density estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vandermeulen, Robert A., Tai, Wai Ming, Aragam, Bryon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912130674982912
author Vandermeulen, Robert A.
Tai, Wai Ming
Aragam, Bryon
author_facet Vandermeulen, Robert A.
Tai, Wai Ming
Aragam, Bryon
contents We show that deep neural networks achieve dimension-independent rates of convergence for learning structured densities such as those arising in image, audio, video, and text applications. More precisely, we demonstrate that neural networks with a simple $L^2$-minimizing loss achieve a rate of $n^{-1/(4+r)}$ in nonparametric density estimation when the underlying density is Markov to a graph whose maximum clique size is at most $r$, and we provide evidence that in the aforementioned applications, this size is typically constant, i.e., $r=O(1)$. We then establish that the optimal rate in $L^1$ is $n^{-1/(2+r)}$ which, compared to the standard nonparametric rate of $n^{-1/(2+d)}$, reveals that the effective dimension of such problems is the size of the largest clique in the Markov random field. These rates are independent of the data's ambient dimension, making them applicable to realistic models of image, sound, video, and text data. Our results provide a novel justification for deep learning's ability to circumvent the curse of dimensionality, demonstrating dimension-independent convergence rates in these contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15095
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dimension-independent rates for structured neural density estimation
Vandermeulen, Robert A.
Tai, Wai Ming
Aragam, Bryon
Machine Learning
Computer Vision and Pattern Recognition
Statistics Theory
62G05, 62G07, 62A09, 62M05, 62M40, 60J10, 60J20
G.3; I.5.1; I.4.10; I.4.m
We show that deep neural networks achieve dimension-independent rates of convergence for learning structured densities such as those arising in image, audio, video, and text applications. More precisely, we demonstrate that neural networks with a simple $L^2$-minimizing loss achieve a rate of $n^{-1/(4+r)}$ in nonparametric density estimation when the underlying density is Markov to a graph whose maximum clique size is at most $r$, and we provide evidence that in the aforementioned applications, this size is typically constant, i.e., $r=O(1)$. We then establish that the optimal rate in $L^1$ is $n^{-1/(2+r)}$ which, compared to the standard nonparametric rate of $n^{-1/(2+d)}$, reveals that the effective dimension of such problems is the size of the largest clique in the Markov random field. These rates are independent of the data's ambient dimension, making them applicable to realistic models of image, sound, video, and text data. Our results provide a novel justification for deep learning's ability to circumvent the curse of dimensionality, demonstrating dimension-independent convergence rates in these contexts.
title Dimension-independent rates for structured neural density estimation
topic Machine Learning
Computer Vision and Pattern Recognition
Statistics Theory
62G05, 62G07, 62A09, 62M05, 62M40, 60J10, 60J20
G.3; I.5.1; I.4.10; I.4.m
url https://arxiv.org/abs/2411.15095