Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruoqi, Luo, Ziwei, Sjölund, Jens, Schön, Thomas B., Mattsson, Per
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916555142463488
author Zhang, Ruoqi
Luo, Ziwei
Sjölund, Jens
Schön, Thomas B.
Mattsson, Per
author_facet Zhang, Ruoqi
Luo, Ziwei
Sjölund, Jens
Schön, Thomas B.
Mattsson, Per
contents This paper presents advanced techniques of training diffusion policies for offline reinforcement learning (RL). At the core is a mean-reverting stochastic differential equation (SDE) that transfers a complex action distribution into a standard Gaussian and then samples actions conditioned on the environment state with a corresponding reverse-time SDE, like a typical diffusion policy. We show that such an SDE has a solution that we can use to calculate the log probability of the policy, yielding an entropy regularizer that improves the exploration of offline datasets. To mitigate the impact of inaccurate value functions from out-of-distribution data points, we further propose to learn the lower confidence bound of Q-ensembles for more robust policy improvement. By combining the entropy-regularized diffusion policy with Q-ensembles in offline RL, our method achieves state-of-the-art performance on most tasks in D4RL benchmarks. Code is available at https://github.com/ruoqizzz/Entropy-Regularized-Diffusion-Policy-with-QEnsemble.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04080
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning
Zhang, Ruoqi
Luo, Ziwei
Sjölund, Jens
Schön, Thomas B.
Mattsson, Per
Machine Learning
Systems and Control
This paper presents advanced techniques of training diffusion policies for offline reinforcement learning (RL). At the core is a mean-reverting stochastic differential equation (SDE) that transfers a complex action distribution into a standard Gaussian and then samples actions conditioned on the environment state with a corresponding reverse-time SDE, like a typical diffusion policy. We show that such an SDE has a solution that we can use to calculate the log probability of the policy, yielding an entropy regularizer that improves the exploration of offline datasets. To mitigate the impact of inaccurate value functions from out-of-distribution data points, we further propose to learn the lower confidence bound of Q-ensembles for more robust policy improvement. By combining the entropy-regularized diffusion policy with Q-ensembles in offline RL, our method achieves state-of-the-art performance on most tasks in D4RL benchmarks. Code is available at https://github.com/ruoqizzz/Entropy-Regularized-Diffusion-Policy-with-QEnsemble.
title Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement Learning
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2402.04080