Goal-Space Planning with Subgoal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lo, Chunlok, Roice, Kevin, Panahi, Parham Mohammad, Jordan, Scott, White, Adam, Mihucz, Gabor, Aminmansour, Farzane, White, Martha
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909120541491200
author Lo, Chunlok
Roice, Kevin
Panahi, Parham Mohammad
Jordan, Scott
White, Adam
Mihucz, Gabor
Aminmansour, Farzane
White, Martha
author_facet Lo, Chunlok
Roice, Kevin
Panahi, Parham Mohammad
Jordan, Scott
White, Adam
Mihucz, Gabor
Aminmansour, Farzane
White, Martha
contents This paper investigates a new approach to model-based reinforcement learning using background planning: mixing (approximate) dynamic programming updates and model-free updates, similar to the Dyna architecture. Background planning with learned models is often worse than model-free alternatives, such as Double DQN, even though the former uses significantly more memory and computation. The fundamental problem is that learned models can be inaccurate and often generate invalid states, especially when iterated many steps. In this paper, we avoid this limitation by constraining background planning to a set of (abstract) subgoals and learning only local, subgoal-conditioned models. This goal-space planning (GSP) approach is more computationally efficient, naturally incorporates temporal abstraction for faster long-horizon planning and avoids learning the transition dynamics entirely. We show that our GSP algorithm can propagate value from an abstract space in a manner that helps a variety of base learners learn significantly faster in different domains.
format Preprint
id arxiv_https___arxiv_org_abs_2206_02902
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Goal-Space Planning with Subgoal Models
Lo, Chunlok
Roice, Kevin
Panahi, Parham Mohammad
Jordan, Scott
White, Adam
Mihucz, Gabor
Aminmansour, Farzane
White, Martha
Machine Learning
Artificial Intelligence
This paper investigates a new approach to model-based reinforcement learning using background planning: mixing (approximate) dynamic programming updates and model-free updates, similar to the Dyna architecture. Background planning with learned models is often worse than model-free alternatives, such as Double DQN, even though the former uses significantly more memory and computation. The fundamental problem is that learned models can be inaccurate and often generate invalid states, especially when iterated many steps. In this paper, we avoid this limitation by constraining background planning to a set of (abstract) subgoals and learning only local, subgoal-conditioned models. This goal-space planning (GSP) approach is more computationally efficient, naturally incorporates temporal abstraction for faster long-horizon planning and avoids learning the transition dynamics entirely. We show that our GSP algorithm can propagate value from an abstract space in a manner that helps a variety of base learners learn significantly faster in different domains.
title Goal-Space Planning with Subgoal Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2206.02902