On the Natural Gradient of the Evidence Lower Bound

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ay, Nihat, van Oostrum, Jesse, Datar, Adwait
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908569802113024
author Ay, Nihat
van Oostrum, Jesse
Datar, Adwait
author_facet Ay, Nihat
van Oostrum, Jesse
Datar, Adwait
contents This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound (ELBO) which plays a central role in generative machine learning. It reveals that the gap between the evidence and its lower bound, the ELBO, has essentially a vanishing natural gradient within unconstrained optimization. As a result, maximization of the ELBO is equivalent to minimization of the Kullback-Leibler divergence from a target distribution, the primary objective function of learning. Building on this insight, we derive a condition under which this equivalence persists even when optimization is constrained to a model. This condition yields a geometric characterization, which we formalize through the notion of a cylindrical model.
format Preprint
id arxiv_https___arxiv_org_abs_2307_11249
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On the Natural Gradient of the Evidence Lower Bound
Ay, Nihat
van Oostrum, Jesse
Datar, Adwait
Machine Learning
Statistics Theory
This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound (ELBO) which plays a central role in generative machine learning. It reveals that the gap between the evidence and its lower bound, the ELBO, has essentially a vanishing natural gradient within unconstrained optimization. As a result, maximization of the ELBO is equivalent to minimization of the Kullback-Leibler divergence from a target distribution, the primary objective function of learning. Building on this insight, we derive a condition under which this equivalence persists even when optimization is constrained to a model. This condition yields a geometric characterization, which we formalize through the notion of a cylindrical model.
title On the Natural Gradient of the Evidence Lower Bound
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2307.11249