Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Jihao Andreas, Antorán, Javier, Padhy, Shreyas, Janz, David, Hernández-Lobato, José Miguel, Terenin, Alexander
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917566930223104
author Lin, Jihao Andreas
Antorán, Javier
Padhy, Shreyas
Janz, David
Hernández-Lobato, José Miguel
Terenin, Alexander
author_facet Lin, Jihao Andreas
Antorán, Javier
Padhy, Shreyas
Janz, David
Hernández-Lobato, José Miguel
Terenin, Alexander
contents Gaussian processes are a powerful framework for quantifying uncertainty and for sequential decision-making but are limited by the requirement of solving linear systems. In general, this has a cubic cost in dataset size and is sensitive to conditioning. We explore stochastic gradient algorithms as a computationally efficient method of approximately solving these linear systems: we develop low-variance optimization objectives for sampling from the posterior and extend these to inducing points. Counterintuitively, stochastic gradient descent often produces accurate predictions, even in cases where it does not converge quickly to the optimum. We explain this through a spectral characterization of the implicit bias from non-convergence. We show that stochastic gradient descent produces predictive distributions close to the true posterior both in regions with sufficient data coverage, and in regions sufficiently far away from the data. Experimentally, stochastic gradient descent achieves state-of-the-art performance on sufficiently large-scale or ill-conditioned regression tasks. Its uncertainty estimates match the performance of significantly more expensive baselines on a large-scale Bayesian optimization task.
format Preprint
id arxiv_https___arxiv_org_abs_2306_11589
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent
Lin, Jihao Andreas
Antorán, Javier
Padhy, Shreyas
Janz, David
Hernández-Lobato, José Miguel
Terenin, Alexander
Machine Learning
Gaussian processes are a powerful framework for quantifying uncertainty and for sequential decision-making but are limited by the requirement of solving linear systems. In general, this has a cubic cost in dataset size and is sensitive to conditioning. We explore stochastic gradient algorithms as a computationally efficient method of approximately solving these linear systems: we develop low-variance optimization objectives for sampling from the posterior and extend these to inducing points. Counterintuitively, stochastic gradient descent often produces accurate predictions, even in cases where it does not converge quickly to the optimum. We explain this through a spectral characterization of the implicit bias from non-convergence. We show that stochastic gradient descent produces predictive distributions close to the true posterior both in regions with sufficient data coverage, and in regions sufficiently far away from the data. Experimentally, stochastic gradient descent achieves state-of-the-art performance on sufficiently large-scale or ill-conditioned regression tasks. Its uncertainty estimates match the performance of significantly more expensive baselines on a large-scale Bayesian optimization task.
title Sampling from Gaussian Process Posteriors using Stochastic Gradient Descent
topic Machine Learning
url https://arxiv.org/abs/2306.11589