Sampling from Energy-based Policies using Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jain, Vineet, Akhound-Sadegh, Tara, Ravanbakhsh, Siamak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915482063339520
author Jain, Vineet
Akhound-Sadegh, Tara
Ravanbakhsh, Siamak
author_facet Jain, Vineet
Akhound-Sadegh, Tara
Ravanbakhsh, Siamak
contents Energy-based policies offer a flexible framework for modeling complex, multimodal behaviors in reinforcement learning (RL). In maximum entropy RL, the optimal policy is a Boltzmann distribution derived from the soft Q-function, but direct sampling from this distribution in continuous action spaces is computationally intractable. As a result, existing methods typically use simpler parametric distributions, like Gaussians, for policy representation -- limiting their ability to capture the full complexity of multimodal action distributions. In this paper, we introduce a diffusion-based approach for sampling from energy-based policies, where the negative Q-function defines the energy function. Based on this approach, we propose an actor-critic method called Diffusion Q-Sampling (DQS) that enables more expressive policy representations, allowing stable learning in diverse environments. We show that our approach enhances sample efficiency in continuous control tasks and captures multimodal behaviors, addressing key limitations of existing methods. Code is available at https://github.com/vineetjain96/Diffusion_Q_Sampling.git
format Preprint
id arxiv_https___arxiv_org_abs_2410_01312
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sampling from Energy-based Policies using Diffusion
Jain, Vineet
Akhound-Sadegh, Tara
Ravanbakhsh, Siamak
Machine Learning
Energy-based policies offer a flexible framework for modeling complex, multimodal behaviors in reinforcement learning (RL). In maximum entropy RL, the optimal policy is a Boltzmann distribution derived from the soft Q-function, but direct sampling from this distribution in continuous action spaces is computationally intractable. As a result, existing methods typically use simpler parametric distributions, like Gaussians, for policy representation -- limiting their ability to capture the full complexity of multimodal action distributions. In this paper, we introduce a diffusion-based approach for sampling from energy-based policies, where the negative Q-function defines the energy function. Based on this approach, we propose an actor-critic method called Diffusion Q-Sampling (DQS) that enables more expressive policy representations, allowing stable learning in diverse environments. We show that our approach enhances sample efficiency in continuous control tasks and captures multimodal behaviors, addressing key limitations of existing methods. Code is available at https://github.com/vineetjain96/Diffusion_Q_Sampling.git
title Sampling from Energy-based Policies using Diffusion
topic Machine Learning
url https://arxiv.org/abs/2410.01312