Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zuyuan, Tang, Sizhe, Lan, Tian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911428409032704
author Zhang, Zuyuan
Tang, Sizhe
Lan, Tian
author_facet Zhang, Zuyuan
Tang, Sizhe
Lan, Tian
contents Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the central pillar of Reinforcement learning (RL) becomes only approximately valid under Non-Markovian. Existing work often focus on practical algorithm designs and offer limited theoretical treatment to address key questions, such as what dynamics are indeed capturable by the Bellman framework and how to inspire new algorithm classes with optimal approximations. In this paper, we present a novel topological viewpoint on temporal-difference (TD) based RL. We show that TD errors can be viewed as 1-cochain in the topological space of state transitions, while Markov dynamics are then interpreted as topological integrability. This novel view enables us to obtain a Hodge-type decomposition of TD errors into an integrable component and a topological residual, through a Bellman-de Rham projection. We further propose HodgeFlow Policy Search (HFPS) by fitting a potential network to minimize the non-integrable projection residual in RL, achieving stability/sensitivity guarantees. In numerical evaluations, HFPS is shown to significantly improve RL performance under non-Markovian.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06939
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
Zhang, Zuyuan
Tang, Sizhe
Lan, Tian
Machine Learning
Artificial Intelligence
Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the central pillar of Reinforcement learning (RL) becomes only approximately valid under Non-Markovian. Existing work often focus on practical algorithm designs and offer limited theoretical treatment to address key questions, such as what dynamics are indeed capturable by the Bellman framework and how to inspire new algorithm classes with optimal approximations. In this paper, we present a novel topological viewpoint on temporal-difference (TD) based RL. We show that TD errors can be viewed as 1-cochain in the topological space of state transitions, while Markov dynamics are then interpreted as topological integrability. This novel view enables us to obtain a Hodge-type decomposition of TD errors into an integrable component and a topological residual, through a Bellman-de Rham projection. We further propose HodgeFlow Policy Search (HFPS) by fitting a potential network to minimize the non-integrable projection residual in RL, achieving stability/sensitivity guarantees. In numerical evaluations, HFPS is shown to significantly improve RL performance under non-Markovian.
title Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.06939