Optimal Transport-Assisted Risk-Sensitive Q-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahrooei, Zahra, Baheri, Ali
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912023724425216
author Shahrooei, Zahra
Baheri, Ali
author_facet Shahrooei, Zahra
Baheri, Ali
contents The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance without considering risk or safety. In contrast, safe reinforcement learning aims to mitigate or avoid unsafe states. This paper presents a risk-sensitive Q-learning algorithm that leverages optimal transport theory to enhance the agent safety. By integrating optimal transport into the Q-learning framework, our approach seeks to optimize the policy's expected return while minimizing the Wasserstein distance between the policy's stationary distribution and a predefined risk distribution, which encapsulates safety preferences from domain experts. We validate the proposed algorithm in a Gridworld environment. The results indicate that our method significantly reduces the frequency of visits to risky states and achieves faster convergence to a stable policy compared to the traditional Q-learning algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11774
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimal Transport-Assisted Risk-Sensitive Q-Learning
Shahrooei, Zahra
Baheri, Ali
Machine Learning
Systems and Control
The primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance without considering risk or safety. In contrast, safe reinforcement learning aims to mitigate or avoid unsafe states. This paper presents a risk-sensitive Q-learning algorithm that leverages optimal transport theory to enhance the agent safety. By integrating optimal transport into the Q-learning framework, our approach seeks to optimize the policy's expected return while minimizing the Wasserstein distance between the policy's stationary distribution and a predefined risk distribution, which encapsulates safety preferences from domain experts. We validate the proposed algorithm in a Gridworld environment. The results indicate that our method significantly reduces the frequency of visits to risky states and achieves faster convergence to a stable policy compared to the traditional Q-learning algorithm.
title Optimal Transport-Assisted Risk-Sensitive Q-Learning
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2406.11774