Reinforcement Learning in POMDP's via Direct Gradient Ascent

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Baxter, Jonathan, Bartlett, Peter L.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915655641464832
author Baxter, Jonathan
Bartlett, Peter L.
author_facet Baxter, Jonathan
Bartlett, Peter L.
contents This paper discusses theoretical and experimental aspects of gradient-based approaches to the direct optimization of policy performance in controlled POMDPs. We introduce GPOMDP, a REINFORCE-like algorithm for estimating an approximation to the gradient of the average reward as a function of the parameters of a stochastic policy. The algorithm's chief advantages are that it requires only a single sample path of the underlying Markov chain, it uses only one free parameter $β\in [0,1)$, which has a natural interpretation in terms of bias-variance trade-off, and it requires no knowledge of the underlying state. We prove convergence of GPOMDP and show how the gradient estimates produced by GPOMDP can be used in a conjugate-gradient procedure to find local optima of the average reward.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02383
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning in POMDP's via Direct Gradient Ascent
Baxter, Jonathan
Bartlett, Peter L.
Machine Learning
This paper discusses theoretical and experimental aspects of gradient-based approaches to the direct optimization of policy performance in controlled POMDPs. We introduce GPOMDP, a REINFORCE-like algorithm for estimating an approximation to the gradient of the average reward as a function of the parameters of a stochastic policy. The algorithm's chief advantages are that it requires only a single sample path of the underlying Markov chain, it uses only one free parameter $β\in [0,1)$, which has a natural interpretation in terms of bias-variance trade-off, and it requires no knowledge of the underlying state. We prove convergence of GPOMDP and show how the gradient estimates produced by GPOMDP can be used in a conjugate-gradient procedure to find local optima of the average reward.
title Reinforcement Learning in POMDP's via Direct Gradient Ascent
topic Machine Learning
url https://arxiv.org/abs/2512.02383