A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914436213637120 |
|---|---|
| author | Calvo-Ordoñez, Sergio Plenk, Jonathan Bergna, Richard Cartea, Alvaro Hernandez-Lobato, Jose Miguel Palla, Konstantina Ciosek, Kamil |
| author_facet | Calvo-Ordoñez, Sergio Plenk, Jonathan Bergna, Richard Cartea, Alvaro Hernandez-Lobato, Jose Miguel Palla, Konstantina Ciosek, Kamil |
| contents | Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific prior mean and with zero observation noise. However, existing formulations have two limitations: (i) the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. To address (ii), we propose a \textit{shifted network} that enables arbitrary prior means and allows obtaining the posterior mean with gradient descent on a single network, without ensembling or kernel inversion. We validate our results with experiments across datasets and architectures, showing that this approach removes key obstacles to the practical use of NTK-GP equivalence in applied Gaussian process modeling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_01556 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks Calvo-Ordoñez, Sergio Plenk, Jonathan Bergna, Richard Cartea, Alvaro Hernandez-Lobato, Jose Miguel Palla, Konstantina Ciosek, Kamil Machine Learning Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific prior mean and with zero observation noise. However, existing formulations have two limitations: (i) the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. To address (ii), we propose a \textit{shifted network} that enables arbitrary prior means and allows obtaining the posterior mean with gradient descent on a single network, without ensembling or kernel inversion. We validate our results with experiments across datasets and architectures, showing that this approach removes key obstacles to the practical use of NTK-GP equivalence in applied Gaussian process modeling. |
| title | A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2502.01556 |