On the Convergence of Black-Box Variational Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Kyurae, Oh, Jisu, Wu, Kaiwen, Ma, Yi-An, Gardner, Jacob R.
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913191637811200
author Kim, Kyurae
Oh, Jisu
Wu, Kaiwen
Ma, Yi-An
Gardner, Jacob R.
author_facet Kim, Kyurae
Oh, Jisu
Wu, Kaiwen
Ma, Yi-An
Gardner, Jacob R.
contents We provide the first convergence guarantee for full black-box variational inference (BBVI), also known as Monte Carlo variational inference. While preliminary investigations worked on simplified versions of BBVI (e.g., bounded domain, bounded support, only optimizing for the scale, and such), our setup does not need any such algorithmic modifications. Our results hold for log-smooth posterior densities with and without strong log-concavity and the location-scale variational family. Also, our analysis reveals that certain algorithm design choices commonly employed in practice, particularly, nonlinear parameterizations of the scale of the variational approximation, can result in suboptimal convergence rates. Fortunately, running BBVI with proximal stochastic gradient descent fixes these limitations, and thus achieves the strongest known convergence rate guarantees. We evaluate this theoretical insight by comparing proximal SGD against other standard implementations of BBVI on large-scale Bayesian inference problems.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15349
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On the Convergence of Black-Box Variational Inference
Kim, Kyurae
Oh, Jisu
Wu, Kaiwen
Ma, Yi-An
Gardner, Jacob R.
Machine Learning
Signal Processing
Optimization and Control
Computation
We provide the first convergence guarantee for full black-box variational inference (BBVI), also known as Monte Carlo variational inference. While preliminary investigations worked on simplified versions of BBVI (e.g., bounded domain, bounded support, only optimizing for the scale, and such), our setup does not need any such algorithmic modifications. Our results hold for log-smooth posterior densities with and without strong log-concavity and the location-scale variational family. Also, our analysis reveals that certain algorithm design choices commonly employed in practice, particularly, nonlinear parameterizations of the scale of the variational approximation, can result in suboptimal convergence rates. Fortunately, running BBVI with proximal stochastic gradient descent fixes these limitations, and thus achieves the strongest known convergence rate guarantees. We evaluate this theoretical insight by comparing proximal SGD against other standard implementations of BBVI on large-scale Bayesian inference problems.
title On the Convergence of Black-Box Variational Inference
topic Machine Learning
Signal Processing
Optimization and Control
Computation
url https://arxiv.org/abs/2305.15349