Wasserstein Barycenter Soft Actor-Critic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahrooei, Zahra, Baheri, Ali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918351834447872
author Shahrooei, Zahra
Baheri, Ali
author_facet Shahrooei, Zahra
Baheri, Ali
contents Deep off-policy actor-critic algorithms have emerged as the leading framework for reinforcement learning in continuous control domains. However, most of these algorithms suffer from poor sample efficiency, especially in environments with sparse rewards. In this paper, we take a step towards addressing this issue by providing a principled directed exploration strategy. We propose Wasserstein Barycenter Soft Actor-Critic (WBSAC) algorithm, which benefits from a pessimistic actor for temporal difference learning and an optimistic actor to promote exploration. This is achieved by using the Wasserstein barycenter of the pessimistic and optimistic policies as the exploration policy and adjusting the degree of exploration throughout the learning process. We compare WBSAC with state-of-the-art off-policy actor-critic algorithms and show that WBSAC is more sample-efficient on MuJoCo continuous control tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Wasserstein Barycenter Soft Actor-Critic
Shahrooei, Zahra
Baheri, Ali
Machine Learning
Systems and Control
Deep off-policy actor-critic algorithms have emerged as the leading framework for reinforcement learning in continuous control domains. However, most of these algorithms suffer from poor sample efficiency, especially in environments with sparse rewards. In this paper, we take a step towards addressing this issue by providing a principled directed exploration strategy. We propose Wasserstein Barycenter Soft Actor-Critic (WBSAC) algorithm, which benefits from a pessimistic actor for temporal difference learning and an optimistic actor to promote exploration. This is achieved by using the Wasserstein barycenter of the pessimistic and optimistic policies as the exploration policy and adjusting the degree of exploration throughout the learning process. We compare WBSAC with state-of-the-art off-policy actor-critic algorithms and show that WBSAC is more sample-efficient on MuJoCo continuous control tasks.
title Wasserstein Barycenter Soft Actor-Critic
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2506.10167