Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yu, Huizhen, Wan, Yi, Sutton, Richard S.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910274162786304
author Yu, Huizhen
Wan, Yi
Sutton, Richard S.
author_facet Yu, Huizhen
Wan, Yi
Sutton, Richard S.
contents This paper investigates the stability and convergence properties of asynchronous stochastic approximation (SA) algorithms, with a focus on extensions relevant to average-reward reinforcement learning. We first extend a stability proof method of Borkar and Meyn to accommodate more general noise conditions than previously considered, thereby yielding broader convergence guarantees for asynchronous SA. To sharpen the convergence analysis, we further examine the shadowing properties of asynchronous SA, building on a dynamical systems approach of Hirsch and Benaïm. These results provide a theoretical foundation for a class of relative value iteration-based reinforcement learning algorithms -- developed and analyzed in a companion paper -- for solving average-reward Markov and semi-Markov decision processes.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03915
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
Yu, Huizhen
Wan, Yi
Sutton, Richard S.
Machine Learning
Optimization and Control
62L20, 90C40, 93E20
This paper investigates the stability and convergence properties of asynchronous stochastic approximation (SA) algorithms, with a focus on extensions relevant to average-reward reinforcement learning. We first extend a stability proof method of Borkar and Meyn to accommodate more general noise conditions than previously considered, thereby yielding broader convergence guarantees for asynchronous SA. To sharpen the convergence analysis, we further examine the shadowing properties of asynchronous SA, building on a dynamical systems approach of Hirsch and Benaïm. These results provide a theoretical foundation for a class of relative value iteration-based reinforcement learning algorithms -- developed and analyzed in a companion paper -- for solving average-reward Markov and semi-Markov decision processes.
title Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
topic Machine Learning
Optimization and Control
62L20, 90C40, 93E20
url https://arxiv.org/abs/2409.03915