A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kaya, Ege C., Hashemi, Abolfazl
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913099616878592
author Kaya, Ege C.
Hashemi, Abolfazl
author_facet Kaya, Ege C.
Hashemi, Abolfazl
contents Recent non-asymptotic analyses have substantially advanced the theory of distributional policy evaluation, but they largely concern synchronous full-state updates under a generative model, model-based estimators, accelerated variants, or different approximation architectures. Standard categorical temporal-difference learning is typically used in a different regime. It asynchronously performs a single-state update at each iteration and, in online settings, is driven by a Markovian trajectory. This leaves an important gap between existing finite-iteration theory and the categorical recursions most closely aligned with practical distributional temporal-difference implementations. We bridge this gap for two categorical policy-evaluation methods: scalar categorical temporal-difference learning in the Cramér geometry and multivariate signed-categorical temporal-difference learning in the maximum mean discrepancy geometry. After suitable isometric embeddings, both algorithms take the form of asynchronous single-state stochastic-approximation recursions that contract in a statewise supremum norm. This permits finite-iteration guarantees in discounted problems under both i.i.d. and Markovian state sampling, and in undiscounted fixed-horizon problems under i.i.d. episodic sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06866
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
Kaya, Ege C.
Hashemi, Abolfazl
Machine Learning
Optimization and Control
Recent non-asymptotic analyses have substantially advanced the theory of distributional policy evaluation, but they largely concern synchronous full-state updates under a generative model, model-based estimators, accelerated variants, or different approximation architectures. Standard categorical temporal-difference learning is typically used in a different regime. It asynchronously performs a single-state update at each iteration and, in online settings, is driven by a Markovian trajectory. This leaves an important gap between existing finite-iteration theory and the categorical recursions most closely aligned with practical distributional temporal-difference implementations. We bridge this gap for two categorical policy-evaluation methods: scalar categorical temporal-difference learning in the Cramér geometry and multivariate signed-categorical temporal-difference learning in the maximum mean discrepancy geometry. After suitable isometric embeddings, both algorithms take the form of asynchronous single-state stochastic-approximation recursions that contract in a statewise supremum norm. This permits finite-iteration guarantees in discounted problems under both i.i.d. and Markovian state sampling, and in undiscounted fixed-horizon problems under i.i.d. episodic sampling.
title A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2605.06866