Reinforced sequential Monte Carlo for amortised sampling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Choi, Sanghyeok, Mittal, Sarthak, Elvira, Víctor, Park, Jinkyoo, Whitammer, Esmeralda S.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913173241593856
author Choi, Sanghyeok
Mittal, Sarthak
Elvira, Víctor
Park, Jinkyoo
Whitammer, Esmeralda S.
author_facet Choi, Sanghyeok
Mittal, Sarthak
Elvira, Víctor
Park, Jinkyoo
Whitammer, Esmeralda S.
contents This paper proposes a synergy of amortised and particle-based methods for sampling from distributions defined by unnormalised density functions. We state a connection between sequential Monte Carlo (SMC) and neural sequential samplers trained by maximum-entropy reinforcement learning (MaxEnt RL), wherein learnt sampling policies and value functions define proposal kernels and twist functions. Exploiting this connection, we introduce an off-policy RL training procedure for the sampler that uses samples from SMC -- using the learnt sampler as a proposal -- as a behaviour policy that better explores the target distribution. We describe techniques for stable joint training of proposals and twist functions and an adaptive weight tempering scheme to reduce training signal variance. Furthermore, building upon past attempts to use experience replay to guide the training of neural samplers, we derive a way to combine historical samples with annealed importance sampling weights within a replay buffer. On synthetic multi-modal targets (in both continuous and discrete spaces) and the Boltzmann distribution of alanine dipeptide conformations, we demonstrate improvements in approximating the true distribution as well as training stability compared to both amortised and Monte Carlo methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforced sequential Monte Carlo for amortised sampling
Choi, Sanghyeok
Mittal, Sarthak
Elvira, Víctor
Park, Jinkyoo
Whitammer, Esmeralda S.
Machine Learning
This paper proposes a synergy of amortised and particle-based methods for sampling from distributions defined by unnormalised density functions. We state a connection between sequential Monte Carlo (SMC) and neural sequential samplers trained by maximum-entropy reinforcement learning (MaxEnt RL), wherein learnt sampling policies and value functions define proposal kernels and twist functions. Exploiting this connection, we introduce an off-policy RL training procedure for the sampler that uses samples from SMC -- using the learnt sampler as a proposal -- as a behaviour policy that better explores the target distribution. We describe techniques for stable joint training of proposals and twist functions and an adaptive weight tempering scheme to reduce training signal variance. Furthermore, building upon past attempts to use experience replay to guide the training of neural samplers, we derive a way to combine historical samples with annealed importance sampling weights within a replay buffer. On synthetic multi-modal targets (in both continuous and discrete spaces) and the Boltzmann distribution of alanine dipeptide conformations, we demonstrate improvements in approximating the true distribution as well as training stability compared to both amortised and Monte Carlo methods.
title Reinforced sequential Monte Carlo for amortised sampling
topic Machine Learning
url https://arxiv.org/abs/2510.11711