Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xiaoshuang, Lin, Yifan, Zhou, Enlu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916981294235648
author Wang, Xiaoshuang
Lin, Yifan
Zhou, Enlu
author_facet Wang, Xiaoshuang
Lin, Yifan
Zhou, Enlu
contents Motivated by many application problems, we consider Markov decision processes (MDPs) with a general loss function and unknown parameters. To mitigate the epistemic uncertainty associated with unknown parameters, we take a Bayesian approach to estimate the parameters from data and impose a coherent risk functional (with respect to the Bayesian posterior distribution) on the loss. Since this formulation usually does not satisfy the interchangeability principle, it does not admit Bellman equations and cannot be solved by approaches based on dynamic programming. Therefore, We propose a policy gradient optimization method, leveraging the dual representation of coherent risk measures and extending the envelope theorem to continuous cases. We then show the stationary analysis of the algorithm with a convergence rate of $\mathcal{O}(T^{-1/2}+r^{-1/2})$, where $T$ is the number of policy gradient iterations and $r$ is the sample size of the gradient estimator. We further extend our algorithm to an episodic setting, and establish the global convergence of the extended algorithm and provide bounds on the number of iterations needed to achieve an error bound $\mathcal{O}(ε)$ in each episode.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
Wang, Xiaoshuang
Lin, Yifan
Zhou, Enlu
Machine Learning
Motivated by many application problems, we consider Markov decision processes (MDPs) with a general loss function and unknown parameters. To mitigate the epistemic uncertainty associated with unknown parameters, we take a Bayesian approach to estimate the parameters from data and impose a coherent risk functional (with respect to the Bayesian posterior distribution) on the loss. Since this formulation usually does not satisfy the interchangeability principle, it does not admit Bellman equations and cannot be solved by approaches based on dynamic programming. Therefore, We propose a policy gradient optimization method, leveraging the dual representation of coherent risk measures and extending the envelope theorem to continuous cases. We then show the stationary analysis of the algorithm with a convergence rate of $\mathcal{O}(T^{-1/2}+r^{-1/2})$, where $T$ is the number of policy gradient iterations and $r$ is the sample size of the gradient estimator. We further extend our algorithm to an episodic setting, and establish the global convergence of the extended algorithm and provide bounds on the number of iterations needed to achieve an error bound $\mathcal{O}(ε)$ in each episode.
title Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
topic Machine Learning
url https://arxiv.org/abs/2509.15509