A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Alver, Safa, Precup, Doina
Format: Preprint
Publié: 2022
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929454416134144
author Alver, Safa
Precup, Doina
author_facet Alver, Safa
Precup, Doina
contents In model-based reinforcement learning (RL), an agent can leverage a learned model to improve its way of behaving in different ways. Two of the prevalent ways to do this are through decision-time and background planning methods. In this study, we are interested in understanding how the value-based versions of these two planning methods will compare against each other across different settings. Towards this goal, we first consider the simplest instantiations of value-based decision-time and background planning methods and provide theoretical results on which one will perform better in the regular RL and transfer learning settings. Then, we consider the modern instantiations of them and provide hypotheses on which one will perform better in the same settings. Finally, we perform illustrative experiments to validate these theoretical results and hypotheses. Overall, our findings suggest that even though value-based versions of the two planning methods perform on par in their simplest instantiations, the modern instantiations of value-based decision-time planning methods can perform on par or better than the modern instantiations of value-based background planning methods in both the regular RL and transfer learning settings.
format Preprint
id arxiv_https___arxiv_org_abs_2206_08442
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
Alver, Safa
Precup, Doina
Machine Learning
Artificial Intelligence
In model-based reinforcement learning (RL), an agent can leverage a learned model to improve its way of behaving in different ways. Two of the prevalent ways to do this are through decision-time and background planning methods. In this study, we are interested in understanding how the value-based versions of these two planning methods will compare against each other across different settings. Towards this goal, we first consider the simplest instantiations of value-based decision-time and background planning methods and provide theoretical results on which one will perform better in the regular RL and transfer learning settings. Then, we consider the modern instantiations of them and provide hypotheses on which one will perform better in the same settings. Finally, we perform illustrative experiments to validate these theoretical results and hypotheses. Overall, our findings suggest that even though value-based versions of the two planning methods perform on par in their simplest instantiations, the modern instantiations of value-based decision-time planning methods can perform on par or better than the modern instantiations of value-based background planning methods in both the regular RL and transfer learning settings.
title A Look at Value-Based Decision-Time vs. Background Planning Methods Across Different Settings
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2206.08442