Accelerated Distributional Temporal Difference Learning with Linear Function Approximation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jin, Kaicheng, Peng, Yang, Yang, Jiansheng, Zhang, Zhihua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914160281911296
author Jin, Kaicheng
Peng, Yang
Yang, Jiansheng
Zhang, Zhihua
author_facet Jin, Kaicheng
Peng, Yang
Yang, Jiansheng
Zhang, Zhihua
contents In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy. Previous works on statistical analysis of distributional TD learning focus mainly on the tabular case. We first consider the linear function approximation setting and conduct a fine-grained analysis of the linear-categorical Bellman equation. Building on this analysis, we further incorporate variance reduction techniques in our new algorithms to establish tight sample complexity bounds independent of the support size $K$ when $K$ is large. Our theoretical results imply that, when employing distributional TD learning with linear function approximation, learning the full distribution of the return function from streaming data is no more difficult than learning its expectation. This work provide new insights into the statistical efficiency of distributional reinforcement learning algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerated Distributional Temporal Difference Learning with Linear Function Approximation
Jin, Kaicheng
Peng, Yang
Yang, Jiansheng
Zhang, Zhihua
Machine Learning
In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD learning is to estimate the return distribution of a discounted Markov decision process for a given policy. Previous works on statistical analysis of distributional TD learning focus mainly on the tabular case. We first consider the linear function approximation setting and conduct a fine-grained analysis of the linear-categorical Bellman equation. Building on this analysis, we further incorporate variance reduction techniques in our new algorithms to establish tight sample complexity bounds independent of the support size $K$ when $K$ is large. Our theoretical results imply that, when employing distributional TD learning with linear function approximation, learning the full distribution of the return function from streaming data is no more difficult than learning its expectation. This work provide new insights into the statistical efficiency of distributional reinforcement learning algorithms.
title Accelerated Distributional Temporal Difference Learning with Linear Function Approximation
topic Machine Learning
url https://arxiv.org/abs/2511.12688