Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ruszczynski, Andrzej, Zhang, Tiangang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913080648138752
author Ruszczynski, Andrzej
Zhang, Tiangang
author_facet Ruszczynski, Andrzej
Zhang, Tiangang
contents For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based $Q$-learning method with multipattern $Q$-factor approximation and we prove a high-probability regret bound of $\mathcal{O}\big(H^2 N^H \sqrt{ K}\big)$, where $H$ is the horizon, $N$ is the mini-batch size, and $K$ is the number of episodes. We also propose an economical version of the $Q$-learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00654
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation
Ruszczynski, Andrzej
Zhang, Tiangang
Machine Learning
Artificial Intelligence
Optimization and Control
90C39, 90C40
I.2.6
For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based $Q$-learning method with multipattern $Q$-factor approximation and we prove a high-probability regret bound of $\mathcal{O}\big(H^2 N^H \sqrt{ K}\big)$, where $H$ is the horizon, $N$ is the mini-batch size, and $K$ is the number of episodes. We also propose an economical version of the $Q$-learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem.
title Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation
topic Machine Learning
Artificial Intelligence
Optimization and Control
90C39, 90C40
I.2.6
url https://arxiv.org/abs/2605.00654