Policy Gradient Methods for Distortion Risk Measures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vijayan, Nithia, A, Prashanth L.
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917581540032512
author Vijayan, Nithia
A, Prashanth L.
author_facet Vijayan, Nithia
A, Prashanth L.
contents We propose policy gradient algorithms which learn risk-sensitive policies in a reinforcement learning (RL) framework. Our proposed algorithms maximize the distortion risk measure (DRM) of the cumulative reward in an episodic Markov decision process in on-policy and off-policy RL settings, respectively. We derive a variant of the policy gradient theorem that caters to the DRM objective, and integrate it with a likelihood ratio-based gradient estimation scheme. We derive non-asymptotic bounds that establish the convergence of our proposed algorithms to an approximate stationary point of the DRM objective.
format Preprint
id arxiv_https___arxiv_org_abs_2107_04422
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Policy Gradient Methods for Distortion Risk Measures
Vijayan, Nithia
A, Prashanth L.
Machine Learning
We propose policy gradient algorithms which learn risk-sensitive policies in a reinforcement learning (RL) framework. Our proposed algorithms maximize the distortion risk measure (DRM) of the cumulative reward in an episodic Markov decision process in on-policy and off-policy RL settings, respectively. We derive a variant of the policy gradient theorem that caters to the DRM objective, and integrate it with a likelihood ratio-based gradient estimation scheme. We derive non-asymptotic bounds that establish the convergence of our proposed algorithms to an approximate stationary point of the DRM objective.
title Policy Gradient Methods for Distortion Risk Measures
topic Machine Learning
url https://arxiv.org/abs/2107.04422