Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Shanyu, Liu, Yang, Yu, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915287722360832
author Han, Shanyu
Liu, Yang
Yu, Xiang
author_facet Han, Shanyu
Liu, Yang
Yu, Xiang
contents We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions
Han, Shanyu
Liu, Yang
Yu, Xiang
Mathematical Finance
Artificial Intelligence
Risk Management
We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.
title Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions
topic Mathematical Finance
Artificial Intelligence
Risk Management
url https://arxiv.org/abs/2505.04553