Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Blaser, Ethan, Wang, Jiuqi, Zhang, Shangtong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917279579504640
author Blaser, Ethan
Wang, Jiuqi
Zhang, Shangtong
author_facet Blaser, Ethan
Wang, Jiuqi
Zhang, Shangtong
contents The average reward is a fundamental performance metric in reinforcement learning (RL) focusing on the long-run performance of an agent. Differential temporal difference (TD) learning algorithms are a major advance for average reward RL as they provide an efficient online method to learn the value functions associated with the average reward in both on-policy and off-policy settings. However, existing convergence guarantees require a local clock in learning rates tied to state visit counts, which practitioners do not use and does not extend beyond tabular settings. We address this limitation by proving the almost sure convergence of on-policy $n$-step differential TD for any $n$ using standard diminishing learning rates without a local clock. We then derive three sufficient conditions under which off-policy $n$-step differential TD also converges without a local clock. These results strengthen the theoretical foundations of differential TD and bring its convergence analysis closer to practical implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16629
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
Blaser, Ethan
Wang, Jiuqi
Zhang, Shangtong
Machine Learning
Artificial Intelligence
The average reward is a fundamental performance metric in reinforcement learning (RL) focusing on the long-run performance of an agent. Differential temporal difference (TD) learning algorithms are a major advance for average reward RL as they provide an efficient online method to learn the value functions associated with the average reward in both on-policy and off-policy settings. However, existing convergence guarantees require a local clock in learning rates tied to state visit counts, which practitioners do not use and does not extend beyond tabular settings. We address this limitation by proving the almost sure convergence of on-policy $n$-step differential TD for any $n$ using standard diminishing learning rates without a local clock. We then derive three sufficient conditions under which off-policy $n$-step differential TD also converges without a local clock. These results strengthen the theoretical foundations of differential TD and bring its convergence analysis closer to practical implementations.
title Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.16629