Recurrent Natural Policy Gradient for POMDPs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cayci, Semih, Eryilmaz, Atilla
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909850997358592
author Cayci, Semih
Eryilmaz, Atilla
author_facet Cayci, Semih
Eryilmaz, Atilla
contents Solving partially observable Markov decision processes (POMDPs) remains a fundamental challenge in reinforcement learning (RL), primarily due to the curse of dimensionality induced by the non-stationarity of optimal policies. In this work, we study a natural actor-critic (NAC) algorithm that integrates recurrent neural network (RNN) architectures into a natural policy gradient (NPG) method and a temporal difference (TD) learning method. This framework leverages the representational capacity of RNNs to address non-stationarity in RL to solve POMDPs while retaining the statistical and computational efficiency of natural gradient methods in RL. We provide non-asymptotic theoretical guarantees for this method, including bounds on sample and iteration complexity to achieve global optimality up to function approximation. Additionally, we characterize pathological cases that stem from long-term dependencies, thereby explaining limitations of RNN-based policy optimization for POMDPs.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18221
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Recurrent Natural Policy Gradient for POMDPs
Cayci, Semih
Eryilmaz, Atilla
Optimization and Control
Machine Learning
Solving partially observable Markov decision processes (POMDPs) remains a fundamental challenge in reinforcement learning (RL), primarily due to the curse of dimensionality induced by the non-stationarity of optimal policies. In this work, we study a natural actor-critic (NAC) algorithm that integrates recurrent neural network (RNN) architectures into a natural policy gradient (NPG) method and a temporal difference (TD) learning method. This framework leverages the representational capacity of RNNs to address non-stationarity in RL to solve POMDPs while retaining the statistical and computational efficiency of natural gradient methods in RL. We provide non-asymptotic theoretical guarantees for this method, including bounds on sample and iteration complexity to achieve global optimality up to function approximation. Additionally, we characterize pathological cases that stem from long-term dependencies, thereby explaining limitations of RNN-based policy optimization for POMDPs.
title Recurrent Natural Policy Gradient for POMDPs
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2405.18221