Is Risk-Sensitive Reinforcement Learning Properly Resolved?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Ruiwen, Liu, Minghuan, Ren, Kan, Luo, Xufang, Zhang, Weinan, Li, Dongsheng
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915590079250432
author Zhou, Ruiwen
Liu, Minghuan
Ren, Kan
Luo, Xufang
Zhang, Weinan
Li, Dongsheng
author_facet Zhou, Ruiwen
Liu, Minghuan
Ren, Kan
Luo, Xufang
Zhang, Weinan
Li, Dongsheng
contents Due to the nature of risk management in learning applicable policies, risk-sensitive reinforcement learning (RSRL) has been realized as an important direction. RSRL is usually achieved by learning risk-sensitive objectives characterized by various risk measures, under the framework of distributional reinforcement learning. However, it remains unclear if the distributional Bellman operator properly optimizes the RSRL objective in the sense of risk measures. In this paper, we prove that the existing RSRL methods do not achieve unbiased optimization and cannot guarantee optimality or even improvements regarding risk measures over accumulated return distributions. To remedy this issue, we further propose a novel algorithm, namely Trajectory Q-Learning (TQL), for RSRL problems with provable policy improvement towards the optimal policy. Based on our new learning architecture, we are free to introduce a general and practical implementation for different risk measures to learn disparate risk-sensitive policies. In the experiments, we verify the learnability of our algorithm and show how our method effectively achieves better performances toward risk-sensitive objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2307_00547
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Is Risk-Sensitive Reinforcement Learning Properly Resolved?
Zhou, Ruiwen
Liu, Minghuan
Ren, Kan
Luo, Xufang
Zhang, Weinan
Li, Dongsheng
Machine Learning
Due to the nature of risk management in learning applicable policies, risk-sensitive reinforcement learning (RSRL) has been realized as an important direction. RSRL is usually achieved by learning risk-sensitive objectives characterized by various risk measures, under the framework of distributional reinforcement learning. However, it remains unclear if the distributional Bellman operator properly optimizes the RSRL objective in the sense of risk measures. In this paper, we prove that the existing RSRL methods do not achieve unbiased optimization and cannot guarantee optimality or even improvements regarding risk measures over accumulated return distributions. To remedy this issue, we further propose a novel algorithm, namely Trajectory Q-Learning (TQL), for RSRL problems with provable policy improvement towards the optimal policy. Based on our new learning architecture, we are free to introduce a general and practical implementation for different risk measures to learn disparate risk-sensitive policies. In the experiments, we verify the learnability of our algorithm and show how our method effectively achieves better performances toward risk-sensitive objectives.
title Is Risk-Sensitive Reinforcement Learning Properly Resolved?
topic Machine Learning
url https://arxiv.org/abs/2307.00547