A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Calvo-Ordoñez, Sergio, Plenk, Jonathan, Bergna, Richard, Cartea, Alvaro, Hernandez-Lobato, Jose Miguel, Palla, Konstantina, Ciosek, Kamil
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914436213637120
author Calvo-Ordoñez, Sergio
Plenk, Jonathan
Bergna, Richard
Cartea, Alvaro
Hernandez-Lobato, Jose Miguel
Palla, Konstantina
Ciosek, Kamil
author_facet Calvo-Ordoñez, Sergio
Plenk, Jonathan
Bergna, Richard
Cartea, Alvaro
Hernandez-Lobato, Jose Miguel
Palla, Konstantina
Ciosek, Kamil
contents Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific prior mean and with zero observation noise. However, existing formulations have two limitations: (i) the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. To address (ii), we propose a \textit{shifted network} that enables arbitrary prior means and allows obtaining the posterior mean with gradient descent on a single network, without ensembling or kernel inversion. We validate our results with experiments across datasets and architectures, showing that this approach removes key obstacles to the practical use of NTK-GP equivalence in applied Gaussian process modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
Calvo-Ordoñez, Sergio
Plenk, Jonathan
Bergna, Richard
Cartea, Alvaro
Hernandez-Lobato, Jose Miguel
Palla, Konstantina
Ciosek, Kamil
Machine Learning
Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific prior mean and with zero observation noise. However, existing formulations have two limitations: (i) the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. To address (ii), we propose a \textit{shifted network} that enables arbitrary prior means and allows obtaining the posterior mean with gradient descent on a single network, without ensembling or kernel inversion. We validate our results with experiments across datasets and architectures, showing that this approach removes key obstacles to the practical use of NTK-GP equivalence in applied Gaussian process modeling.
title A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
topic Machine Learning
url https://arxiv.org/abs/2502.01556