Implicit Updates for Average-Reward Temporal Difference Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Hwanwoo, Cho, Dongkyu Derek, Laber, Eric
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908579674456064
author Kim, Hwanwoo
Cho, Dongkyu Derek
Laber, Eric
author_facet Kim, Hwanwoo
Cho, Dongkyu Derek
Laber, Eric
contents Temporal difference (TD) learning is a cornerstone of reinforcement learning. In the average-reward setting, standard TD($λ$) is highly sensitive to the choice of step-size and thus requires careful tuning to maintain numerical stability. We introduce average-reward implicit TD($λ$), which employs an implicit fixed point update to provide data-adaptive stabilization while preserving the per iteration computational complexity of standard average-reward TD($λ$). In contrast to prior finite-time analyses of average-reward TD($λ$), which impose restrictive step-size conditions, we establish finite-time error bounds for the implicit variant under substantially weaker step-size requirements. Empirically, average-reward implicit TD($λ$) operates reliably over a much broader range of step-sizes and exhibits markedly improved numerical stability. This enables more efficient policy evaluation and policy learning, highlighting its effectiveness as a robust alternative to average-reward TD($λ$).
format Preprint
id arxiv_https___arxiv_org_abs_2510_06149
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Implicit Updates for Average-Reward Temporal Difference Learning
Kim, Hwanwoo
Cho, Dongkyu Derek
Laber, Eric
Machine Learning
Temporal difference (TD) learning is a cornerstone of reinforcement learning. In the average-reward setting, standard TD($λ$) is highly sensitive to the choice of step-size and thus requires careful tuning to maintain numerical stability. We introduce average-reward implicit TD($λ$), which employs an implicit fixed point update to provide data-adaptive stabilization while preserving the per iteration computational complexity of standard average-reward TD($λ$). In contrast to prior finite-time analyses of average-reward TD($λ$), which impose restrictive step-size conditions, we establish finite-time error bounds for the implicit variant under substantially weaker step-size requirements. Empirically, average-reward implicit TD($λ$) operates reliably over a much broader range of step-sizes and exhibits markedly improved numerical stability. This enables more efficient policy evaluation and policy learning, highlighting its effectiveness as a robust alternative to average-reward TD($λ$).
title Implicit Updates for Average-Reward Temporal Difference Learning
topic Machine Learning
url https://arxiv.org/abs/2510.06149