Saved in:
Bibliographic Details
Main Authors: Messaoud, Safa, Mokeddem, Billel, Xue, Zhenghai, Pang, Linsey, An, Bo, Chen, Haipeng, Chawla, Sanjay
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2405.00987
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910431595986944
author Messaoud, Safa
Mokeddem, Billel
Xue, Zhenghai
Pang, Linsey
An, Bo
Chen, Haipeng
Chawla, Sanjay
author_facet Messaoud, Safa
Mokeddem, Billel
Xue, Zhenghai
Pang, Linsey
An, Bo
Chen, Haipeng
Chawla, Sanjay
contents Learning expressive stochastic policies instead of deterministic ones has been proposed to achieve better stability, sample complexity, and robustness. Notably, in Maximum Entropy Reinforcement Learning (MaxEnt RL), the policy is modeled as an expressive Energy-Based Model (EBM) over the Q-values. However, this formulation requires the estimation of the entropy of such EBMs, which is an open problem. To address this, previous MaxEnt RL methods either implicitly estimate the entropy, resulting in high computational complexity and variance (SQL), or follow a variational inference procedure that fits simplified actor distributions (e.g., Gaussian) for tractability (SAC). We propose Stein Soft Actor-Critic (S$^2$AC), a MaxEnt RL algorithm that learns expressive policies without compromising efficiency. Specifically, S$^2$AC uses parameterized Stein Variational Gradient Descent (SVGD) as the underlying policy. We derive a closed-form expression of the entropy of such policies. Our formula is computationally efficient and only depends on first-order derivatives and vector products. Empirical results show that S$^2$AC yields more optimal solutions to the MaxEnt objective than SQL and SAC in the multi-goal environment, and outperforms SAC and SQL on the MuJoCo benchmark. Our code is available at: https://github.com/SafaMessaoud/S2AC-Energy-Based-RL-with-Stein-Soft-Actor-Critic
format Preprint
id arxiv_https___arxiv_org_abs_2405_00987
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle S$^2$AC: Energy-Based Reinforcement Learning with Stein Soft Actor Critic
Messaoud, Safa
Mokeddem, Billel
Xue, Zhenghai
Pang, Linsey
An, Bo
Chen, Haipeng
Chawla, Sanjay
Machine Learning
Learning expressive stochastic policies instead of deterministic ones has been proposed to achieve better stability, sample complexity, and robustness. Notably, in Maximum Entropy Reinforcement Learning (MaxEnt RL), the policy is modeled as an expressive Energy-Based Model (EBM) over the Q-values. However, this formulation requires the estimation of the entropy of such EBMs, which is an open problem. To address this, previous MaxEnt RL methods either implicitly estimate the entropy, resulting in high computational complexity and variance (SQL), or follow a variational inference procedure that fits simplified actor distributions (e.g., Gaussian) for tractability (SAC). We propose Stein Soft Actor-Critic (S$^2$AC), a MaxEnt RL algorithm that learns expressive policies without compromising efficiency. Specifically, S$^2$AC uses parameterized Stein Variational Gradient Descent (SVGD) as the underlying policy. We derive a closed-form expression of the entropy of such policies. Our formula is computationally efficient and only depends on first-order derivatives and vector products. Empirical results show that S$^2$AC yields more optimal solutions to the MaxEnt objective than SQL and SAC in the multi-goal environment, and outperforms SAC and SQL on the MuJoCo benchmark. Our code is available at: https://github.com/SafaMessaoud/S2AC-Energy-Based-RL-with-Stein-Soft-Actor-Critic
title S$^2$AC: Energy-Based Reinforcement Learning with Stein Soft Actor Critic
topic Machine Learning
url https://arxiv.org/abs/2405.00987