Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Morimura, Tetsuro, Ota, Kazuhiro, Abe, Kenshi, Zhang, Peinan
Format: Preprint
Publié: 2022
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913416575188992
author Morimura, Tetsuro
Ota, Kazuhiro
Abe, Kenshi
Zhang, Peinan
author_facet Morimura, Tetsuro
Ota, Kazuhiro
Abe, Kenshi
Zhang, Peinan
contents Policy gradient (PG) is a reinforcement learning (RL) approach that optimizes a parameterized policy model for an expected return using gradient ascent. While PG can work well even in non-Markovian environments, it may encounter plateaus or peakiness issues. As another successful RL approach, algorithms based on Monte Carlo Tree Search (MCTS), which include AlphaZero, have obtained groundbreaking results, especially in the game-playing domain. They are also effective when applied to non-Markov decision processes. However, the standard MCTS is a method for decision-time planning, which differs from the online RL setting. In this work, we first introduce Monte Carlo Tree Learning (MCTL), an adaptation of MCTS for online RL setups. We then explore a combined policy approach of PG and MCTL to leverage their strengths. We derive conditions for asymptotic convergence with the results of a two-timescale stochastic approximation and propose an algorithm that satisfies these conditions and converges to a reasonable solution. Our numerical experiments validate the effectiveness of the proposed methods.
format Preprint
id arxiv_https___arxiv_org_abs_2206_01011
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
Morimura, Tetsuro
Ota, Kazuhiro
Abe, Kenshi
Zhang, Peinan
Machine Learning
Artificial Intelligence
Policy gradient (PG) is a reinforcement learning (RL) approach that optimizes a parameterized policy model for an expected return using gradient ascent. While PG can work well even in non-Markovian environments, it may encounter plateaus or peakiness issues. As another successful RL approach, algorithms based on Monte Carlo Tree Search (MCTS), which include AlphaZero, have obtained groundbreaking results, especially in the game-playing domain. They are also effective when applied to non-Markov decision processes. However, the standard MCTS is a method for decision-time planning, which differs from the online RL setting. In this work, we first introduce Monte Carlo Tree Learning (MCTL), an adaptation of MCTS for online RL setups. We then explore a combined policy approach of PG and MCTL to leverage their strengths. We derive conditions for asymptotic convergence with the results of a two-timescale stochastic approximation and propose an algorithm that satisfies these conditions and converges to a reasonable solution. Our numerical experiments validate the effectiveness of the proposed methods.
title Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2206.01011