Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Dake, Lyu, Boxiang, Qiu, Shuang, Kolar, Mladen, Zhang, Tong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911951051816960
author Zhang, Dake
Lyu, Boxiang
Qiu, Shuang
Kolar, Mladen
Zhang, Tong
author_facet Zhang, Dake
Lyu, Boxiang
Qiu, Shuang
Kolar, Mladen
Zhang, Tong
contents We study risk-sensitive reinforcement learning (RL), a crucial field due to its ability to enhance decision-making in scenarios where it is essential to manage uncertainty and minimize potential adverse outcomes. Particularly, our work focuses on applying the entropic risk measure to RL problems. While existing literature primarily investigates the online setting, there remains a large gap in understanding how to efficiently derive a near-optimal policy based on this risk measure using only a pre-collected dataset. We center on the linear Markov Decision Process (MDP) setting, a well-regarded theoretical framework that has yet to be examined from a risk-sensitive standpoint. In response, we introduce two provably sample-efficient algorithms. We begin by presenting a risk-sensitive pessimistic value iteration algorithm, offering a tight analysis by leveraging the structure of the risk-sensitive performance measure. To further improve the obtained bounds, we propose another pessimistic algorithm that utilizes variance information and reference-advantage decomposition, effectively improving both the dependence on the space dimension $d$ and the risk-sensitivity factor. To the best of our knowledge, we obtain the first provably efficient risk-sensitive offline RL algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2407_07631
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
Zhang, Dake
Lyu, Boxiang
Qiu, Shuang
Kolar, Mladen
Zhang, Tong
Machine Learning
Optimization and Control
Statistics Theory
We study risk-sensitive reinforcement learning (RL), a crucial field due to its ability to enhance decision-making in scenarios where it is essential to manage uncertainty and minimize potential adverse outcomes. Particularly, our work focuses on applying the entropic risk measure to RL problems. While existing literature primarily investigates the online setting, there remains a large gap in understanding how to efficiently derive a near-optimal policy based on this risk measure using only a pre-collected dataset. We center on the linear Markov Decision Process (MDP) setting, a well-regarded theoretical framework that has yet to be examined from a risk-sensitive standpoint. In response, we introduce two provably sample-efficient algorithms. We begin by presenting a risk-sensitive pessimistic value iteration algorithm, offering a tight analysis by leveraging the structure of the risk-sensitive performance measure. To further improve the obtained bounds, we propose another pessimistic algorithm that utilizes variance information and reference-advantage decomposition, effectively improving both the dependence on the space dimension $d$ and the risk-sensitivity factor. To the best of our knowledge, we obtain the first provably efficient risk-sensitive offline RL algorithms.
title Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
topic Machine Learning
Optimization and Control
Statistics Theory
url https://arxiv.org/abs/2407.07631