Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baheri, Ali, Shahrooei, Zahra, Salgarkar, Chirayu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913724287156224
author Baheri, Ali
Shahrooei, Zahra
Salgarkar, Chirayu
author_facet Baheri, Ali
Shahrooei, Zahra
Salgarkar, Chirayu
contents We present Wasserstein Adaptive Value Estimation for Actor-Critic (WAVE), an approach to enhance stability in deep reinforcement learning through adaptive Wasserstein regularization. Our method addresses the inherent instability of actor-critic algorithms by incorporating an adaptively weighted Wasserstein regularization term into the critic's loss function. We prove that WAVE achieves $\mathcal{O}\left(\frac{1}{k}\right)$ convergence rate for the critic's mean squared error and provide theoretical guarantees for stability through Wasserstein-based regularization. Using the Sinkhorn approximation for computational efficiency, our approach automatically adjusts the regularization based on the agent's performance. Theoretical analysis and experimental results demonstrate that WAVE achieves superior performance compared to standard actor-critic methods.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning
Baheri, Ali
Shahrooei, Zahra
Salgarkar, Chirayu
Machine Learning
Systems and Control
We present Wasserstein Adaptive Value Estimation for Actor-Critic (WAVE), an approach to enhance stability in deep reinforcement learning through adaptive Wasserstein regularization. Our method addresses the inherent instability of actor-critic algorithms by incorporating an adaptively weighted Wasserstein regularization term into the critic's loss function. We prove that WAVE achieves $\mathcal{O}\left(\frac{1}{k}\right)$ convergence rate for the critic's mean squared error and provide theoretical guarantees for stability through Wasserstein-based regularization. Using the Sinkhorn approximation for computational efficiency, our approach automatically adjusts the regularization based on the agent's performance. Theoretical analysis and experimental results demonstrate that WAVE achieves superior performance compared to standard actor-critic methods.
title Wasserstein Adaptive Value Estimation for Actor-Critic Reinforcement Learning
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2501.10605