Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Patel, Bhrij, Suttle, Wesley A., Koppel, Alec, Aggarwal, Vaneet, Sadler, Brian M., Bedi, Amrit Singh, Manocha, Dinesh
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909228545867776
author Patel, Bhrij
Suttle, Wesley A.
Koppel, Alec
Aggarwal, Vaneet
Sadler, Brian M.
Bedi, Amrit Singh
Manocha, Dinesh
author_facet Patel, Bhrij
Suttle, Wesley A.
Koppel, Alec
Aggarwal, Vaneet
Sadler, Brian M.
Bedi, Amrit Singh
Manocha, Dinesh
contents In the context of average-reward reinforcement learning, the requirement for oracle knowledge of the mixing time, a measure of the duration a Markov chain under a fixed policy needs to achieve its stationary distribution, poses a significant challenge for the global convergence of policy gradient methods. This requirement is particularly problematic due to the difficulty and expense of estimating mixing time in environments with large state spaces, leading to the necessity of impractically long trajectories for effective gradient estimation in practical applications. To address this limitation, we consider the Multi-level Actor-Critic (MAC) framework, which incorporates a Multi-level Monte-Carlo (MLMC) gradient estimator. With our approach, we effectively alleviate the dependency on mixing time knowledge, a first for average-reward MDPs global convergence. Furthermore, our approach exhibits the tightest available dependence of $\mathcal{O}\left( \sqrt{τ_{mix}} \right)$known from prior work. With a 2D grid world goal-reaching navigation experiment, we demonstrate that MAC outperforms the existing state-of-the-art policy gradient-based method for average reward settings.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11925
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
Patel, Bhrij
Suttle, Wesley A.
Koppel, Alec
Aggarwal, Vaneet
Sadler, Brian M.
Bedi, Amrit Singh
Manocha, Dinesh
Machine Learning
In the context of average-reward reinforcement learning, the requirement for oracle knowledge of the mixing time, a measure of the duration a Markov chain under a fixed policy needs to achieve its stationary distribution, poses a significant challenge for the global convergence of policy gradient methods. This requirement is particularly problematic due to the difficulty and expense of estimating mixing time in environments with large state spaces, leading to the necessity of impractically long trajectories for effective gradient estimation in practical applications. To address this limitation, we consider the Multi-level Actor-Critic (MAC) framework, which incorporates a Multi-level Monte-Carlo (MLMC) gradient estimator. With our approach, we effectively alleviate the dependency on mixing time knowledge, a first for average-reward MDPs global convergence. Furthermore, our approach exhibits the tightest available dependence of $\mathcal{O}\left( \sqrt{τ_{mix}} \right)$known from prior work. With a 2D grid world goal-reaching navigation experiment, we demonstrate that MAC outperforms the existing state-of-the-art policy gradient-based method for average reward settings.
title Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
topic Machine Learning
url https://arxiv.org/abs/2403.11925