Reinforcement Learning for Adaptive MCMC

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Congye, Chen, Wilson, Kanagawa, Heishiro, Oates, Chris. J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910457369985024
author Wang, Congye
Chen, Wilson
Kanagawa, Heishiro
Oates, Chris. J.
author_facet Wang, Congye
Chen, Wilson
Kanagawa, Heishiro
Oates, Chris. J.
contents An informal observation, made by several authors, is that the adaptive design of a Markov transition kernel has the flavour of a reinforcement learning task. Yet, to-date it has remained unclear how to actually exploit modern reinforcement learning technologies for adaptive MCMC. The aim of this paper is to set out a general framework, called Reinforcement Learning Metropolis--Hastings, that is theoretically supported and empirically validated. Our principal focus is on learning fast-mixing Metropolis--Hastings transition kernels, which we cast as deterministic policies and optimise via a policy gradient. Control of the learning rate provably ensures conditions for ergodicity are satisfied. The methodology is used to construct a gradient-free sampler that out-performs a popular gradient-free adaptive Metropolis--Hastings algorithm on $\approx 90 \%$ of tasks in the PosteriorDB benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13574
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reinforcement Learning for Adaptive MCMC
Wang, Congye
Chen, Wilson
Kanagawa, Heishiro
Oates, Chris. J.
Computation
Machine Learning
An informal observation, made by several authors, is that the adaptive design of a Markov transition kernel has the flavour of a reinforcement learning task. Yet, to-date it has remained unclear how to actually exploit modern reinforcement learning technologies for adaptive MCMC. The aim of this paper is to set out a general framework, called Reinforcement Learning Metropolis--Hastings, that is theoretically supported and empirically validated. Our principal focus is on learning fast-mixing Metropolis--Hastings transition kernels, which we cast as deterministic policies and optimise via a policy gradient. Control of the learning rate provably ensures conditions for ergodicity are satisfied. The methodology is used to construct a gradient-free sampler that out-performs a popular gradient-free adaptive Metropolis--Hastings algorithm on $\approx 90 \%$ of tasks in the PosteriorDB benchmark.
title Reinforcement Learning for Adaptive MCMC
topic Computation
Machine Learning
url https://arxiv.org/abs/2405.13574