Posterior and variational inference for deep neural networks with heavy-tailed weights

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Castillo, Ismaël, Egels, Paul
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912327464386560
author Castillo, Ismaël
Egels, Paul
author_facet Castillo, Ismaël
Egels, Paul
contents We consider deep neural networks in a Bayesian framework with a prior distribution sampling the network weights at random. Following a recent idea of Agapiou and Castillo (2023), who show that heavy-tailed prior distributions achieve automatic adaptation to smoothness, we introduce a simple Bayesian deep learning prior based on heavy-tailed weights and ReLU activation. We show that the corresponding posterior distribution achieves near-optimal minimax contraction rates, simultaneously adaptive to both intrinsic dimension and smoothness of the underlying function, in a variety of contexts including nonparametric regression, geometric data and Besov spaces. While most works so far need a form of model selection built-in within the prior distribution, a key aspect of our approach is that it does not require to sample hyperparameters to learn the architecture of the network. We also provide variational Bayes counterparts of the results, that show that mean-field variational approximations still benefit from near-optimal theoretical support.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03369
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Posterior and variational inference for deep neural networks with heavy-tailed weights
Castillo, Ismaël
Egels, Paul
Machine Learning
Statistics Theory
We consider deep neural networks in a Bayesian framework with a prior distribution sampling the network weights at random. Following a recent idea of Agapiou and Castillo (2023), who show that heavy-tailed prior distributions achieve automatic adaptation to smoothness, we introduce a simple Bayesian deep learning prior based on heavy-tailed weights and ReLU activation. We show that the corresponding posterior distribution achieves near-optimal minimax contraction rates, simultaneously adaptive to both intrinsic dimension and smoothness of the underlying function, in a variety of contexts including nonparametric regression, geometric data and Besov spaces. While most works so far need a form of model selection built-in within the prior distribution, a key aspect of our approach is that it does not require to sample hyperparameters to learn the architecture of the network. We also provide variational Bayes counterparts of the results, that show that mean-field variational approximations still benefit from near-optimal theoretical support.
title Posterior and variational inference for deep neural networks with heavy-tailed weights
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2406.03369