Large deviation principles for convolutional Bayesian neural networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bassetti, Federico, De Palma, Vassili, Ladelli, Lucia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908870091210752
author Bassetti, Federico
De Palma, Vassili
Ladelli, Lucia
author_facet Bassetti, Federico
De Palma, Vassili
Ladelli, Lucia
contents While suitably scaled CNNs with Gaussian initialization are known to converge to Gaussian processes as the number of channels diverges, little is known beyond this Gaussian limit. We establish a large deviation principle (LDP) for convolutional neural networks in the infinite-channel regime. We consider a broad class of multidimensional CNN architectures characterized by general receptive fields encoded through a patch-extractor function satisfying mild structural assumptions. Our main result establishes a large deviation principle (LDP) for the sequence of conditional covariance matrices under Gaussian prior distribution on the weights. We further derive an LDP for the posterior distribution obtained by conditioning on a finite number of observations. In addition, we provide a streamlined proof of the concentration of the conditional covariances and of the Gaussian equivalence of the network. To the best of our knowledge, this is the first large deviation principle established for convolutional neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06023
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Large deviation principles for convolutional Bayesian neural networks
Bassetti, Federico
De Palma, Vassili
Ladelli, Lucia
Probability
Machine Learning
While suitably scaled CNNs with Gaussian initialization are known to converge to Gaussian processes as the number of channels diverges, little is known beyond this Gaussian limit. We establish a large deviation principle (LDP) for convolutional neural networks in the infinite-channel regime. We consider a broad class of multidimensional CNN architectures characterized by general receptive fields encoded through a patch-extractor function satisfying mild structural assumptions. Our main result establishes a large deviation principle (LDP) for the sequence of conditional covariance matrices under Gaussian prior distribution on the weights. We further derive an LDP for the posterior distribution obtained by conditioning on a finite number of observations. In addition, we provide a streamlined proof of the concentration of the conditional covariances and of the Gaussian equivalence of the network. To the best of our knowledge, this is the first large deviation principle established for convolutional neural networks.
title Large deviation principles for convolutional Bayesian neural networks
topic Probability
Machine Learning
url https://arxiv.org/abs/2603.06023