Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Huizhen, Wan, Yi, Sutton, Richard S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917129384624128
author Yu, Huizhen
Wan, Yi
Sutton, Richard S.
author_facet Yu, Huizhen
Wan, Yi
Sutton, Richard S.
contents This paper applies the authors' recent results on asynchronous stochastic approximation (SA) in the Borkar-Meyn framework to reinforcement learning in average-reward semi-Markov decision processes (SMDPs). We establish the convergence of an asynchronous SA analogue of Schweitzer's classical relative value iteration algorithm, RVI Q-learning, for finite-space, weakly communicating SMDPs. In particular, we show that the algorithm converges almost surely to a compact, connected subset of solutions to the average-reward optimality equation, with convergence to a unique, sample path-dependent solution under additional stepsize and asynchrony conditions. Moreover, to make full use of the SA framework, we introduce new monotonicity conditions for estimating the optimal reward rate in RVI Q-learning. These conditions substantially expand the previously considered algorithmic framework and are addressed through novel arguments in the stability and convergence analysis of RVI Q-learning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
Yu, Huizhen
Wan, Yi
Sutton, Richard S.
Machine Learning
Optimization and Control
90C40, 62L20, 93E20
This paper applies the authors' recent results on asynchronous stochastic approximation (SA) in the Borkar-Meyn framework to reinforcement learning in average-reward semi-Markov decision processes (SMDPs). We establish the convergence of an asynchronous SA analogue of Schweitzer's classical relative value iteration algorithm, RVI Q-learning, for finite-space, weakly communicating SMDPs. In particular, we show that the algorithm converges almost surely to a compact, connected subset of solutions to the average-reward optimality equation, with convergence to a unique, sample path-dependent solution under additional stepsize and asynchrony conditions. Moreover, to make full use of the SA framework, we introduce new monotonicity conditions for estimating the optimal reward rate in RVI Q-learning. These conditions substantially expand the previously considered algorithmic framework and are addressed through novel arguments in the stability and convergence analysis of RVI Q-learning.
title Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
topic Machine Learning
Optimization and Control
90C40, 62L20, 93E20
url https://arxiv.org/abs/2512.06218