Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation
Fuente:
arXiv
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913080648138752 |
|---|---|
| author | Ruszczynski, Andrzej Zhang, Tiangang |
| author_facet | Ruszczynski, Andrzej Zhang, Tiangang |
| contents | For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based $Q$-learning method with multipattern $Q$-factor approximation and we prove a high-probability regret bound of $\mathcal{O}\big(H^2 N^H \sqrt{ K}\big)$, where $H$ is the horizon, $N$ is the mini-batch size, and $K$ is the number of episodes. We also propose an economical version of the $Q$-learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_00654 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation Ruszczynski, Andrzej Zhang, Tiangang Machine Learning Artificial Intelligence Optimization and Control 90C39, 90C40 I.2.6 For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based $Q$-learning method with multipattern $Q$-factor approximation and we prove a high-probability regret bound of $\mathcal{O}\big(H^2 N^H \sqrt{ K}\big)$, where $H$ is the horizon, $N$ is the mini-batch size, and $K$ is the number of episodes. We also propose an economical version of the $Q$-learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem. |
| title | Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation |
| topic | Machine Learning Artificial Intelligence Optimization and Control 90C39, 90C40 I.2.6 |
| url | https://arxiv.org/abs/2605.00654 |