Policy Gradient Methods for Non-Markovian Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kar, Avik, Chandak, Siddharth, Singh, Rahul, Sinhahajari, Soumitra, Moulines, Eric, Bhatnagar, Shalabh, Bambos, Nicholas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911671855874048
author Kar, Avik
Chandak, Siddharth
Singh, Rahul
Sinhahajari, Soumitra
Moulines, Eric
Bhatnagar, Shalabh
Bambos, Nicholas
author_facet Kar, Avik
Chandak, Siddharth
Singh, Rahul
Sinhahajari, Soumitra
Moulines, Eric
Bhatnagar, Shalabh
Bambos, Nicholas
contents We study policy gradient methods for reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards depend on the entire interaction history. To handle this dependence, the agent maintains an internal state that is recursively updated to provide a compact summary of past observations and actions. In contrast to approaches that treat the agent state dynamics as fixed or learn it via predictive objectives, we propose a reward-centric formulation that jointly optimizes the agent state dynamics and the control policy to maximize the expected cumulative reward. To this end, we consider a class of Agent State-Markov (ASM) policies, comprising an agent state dynamics and a control policy that maps the agent state to actions. We establish a novel policy gradient theorem for ASM policies, extending the classical policy gradient results from the Markovian setting to episodic and infinite-horizon discounted NMDPs. Building on this gradient expression, we propose the Agent State-Markov Policy Gradient (ASMPG) algorithm, which leverages the recursive structure of the agent state dynamics for efficient optimization. We establish finite-time and almost sure convergence guarantees, and empirically demonstrate that, on a range of non-Markovian tasks, ASMPG outperforms baselines that learn state representations via predictive objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10816
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Policy Gradient Methods for Non-Markovian Reinforcement Learning
Kar, Avik
Chandak, Siddharth
Singh, Rahul
Sinhahajari, Soumitra
Moulines, Eric
Bhatnagar, Shalabh
Bambos, Nicholas
Machine Learning
Artificial Intelligence
We study policy gradient methods for reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards depend on the entire interaction history. To handle this dependence, the agent maintains an internal state that is recursively updated to provide a compact summary of past observations and actions. In contrast to approaches that treat the agent state dynamics as fixed or learn it via predictive objectives, we propose a reward-centric formulation that jointly optimizes the agent state dynamics and the control policy to maximize the expected cumulative reward. To this end, we consider a class of Agent State-Markov (ASM) policies, comprising an agent state dynamics and a control policy that maps the agent state to actions. We establish a novel policy gradient theorem for ASM policies, extending the classical policy gradient results from the Markovian setting to episodic and infinite-horizon discounted NMDPs. Building on this gradient expression, we propose the Agent State-Markov Policy Gradient (ASMPG) algorithm, which leverages the recursive structure of the agent state dynamics for efficient optimization. We establish finite-time and almost sure convergence guarantees, and empirically demonstrate that, on a range of non-Markovian tasks, ASMPG outperforms baselines that learn state representations via predictive objectives.
title Policy Gradient Methods for Non-Markovian Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.10816