Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bassetti, Federico, Gherardi, Marco, Ingrosso, Alessandro, Pastore, Mauro, Rotondo, Pietro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909659371143168
author Bassetti, Federico
Gherardi, Marco
Ingrosso, Alessandro
Pastore, Mauro
Rotondo, Pietro
author_facet Bassetti, Federico
Gherardi, Marco
Ingrosso, Alessandro
Pastore, Mauro
Rotondo, Pietro
contents Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript, we provide rigorous results for the statistics of functions implemented by the aforementioned class of networks, thus moving closer to a complete characterization of feature learning in the Bayesian setting. Our results include: (i) an exact and elementary non-asymptotic integral representation for the joint prior distribution over the outputs, given in terms of a mixture of Gaussians; (ii) an analytical formula for the posterior distribution in the case of squared error loss function (Gaussian likelihood); (iii) a quantitative description of the feature learning infinite-width regime, using large deviation theory. From a physical perspective, deep architectures with multiple outputs or convolutional layers represent different manifestations of kernel shape renormalization, and our work provides a dictionary that translates this physics intuition and terminology into rigorous Bayesian statistics.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03260
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers
Bassetti, Federico
Gherardi, Marco
Ingrosso, Alessandro
Pastore, Mauro
Rotondo, Pietro
Machine Learning
Disordered Systems and Neural Networks
Statistics Theory
62E20, 62E15, 82B44
Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript, we provide rigorous results for the statistics of functions implemented by the aforementioned class of networks, thus moving closer to a complete characterization of feature learning in the Bayesian setting. Our results include: (i) an exact and elementary non-asymptotic integral representation for the joint prior distribution over the outputs, given in terms of a mixture of Gaussians; (ii) an analytical formula for the posterior distribution in the case of squared error loss function (Gaussian likelihood); (iii) a quantitative description of the feature learning infinite-width regime, using large deviation theory. From a physical perspective, deep architectures with multiple outputs or convolutional layers represent different manifestations of kernel shape renormalization, and our work provides a dictionary that translates this physics intuition and terminology into rigorous Bayesian statistics.
title Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers
topic Machine Learning
Disordered Systems and Neural Networks
Statistics Theory
62E20, 62E15, 82B44
url https://arxiv.org/abs/2406.03260