Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Aminmansour, Farzane, Jafferjee, Taher, Imani, Ehsan, Talvitie, Erin, Bowling, Micheal, White, Martha
Format: Preprint
Veröffentlicht: 2020
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908935071465472
author Aminmansour, Farzane
Jafferjee, Taher
Imani, Ehsan
Talvitie, Erin
Bowling, Micheal
White, Martha
author_facet Aminmansour, Farzane
Jafferjee, Taher
Imani, Ehsan
Talvitie, Erin
Bowling, Micheal
White, Martha
contents Dyna-style reinforcement learning (RL) agents improve sample efficiency over model-free RL agents by updating the value function with simulated experience generated by an environment model. However, it is often difficult to learn accurate models of environment dynamics, and even small errors may result in failure of Dyna agents. In this paper, we highlight that one potential cause of that failure is bootstrapping off of the values of simulated states, and introduce a new Dyna algorithm to avoid this failure. We discuss a design space of Dyna algorithms, based on using successor or predecessor models -- simulating forwards or backwards -- and using one-step or multi-step updates. Three of the variants have been explored, but surprisingly the fourth variant has not: using predecessor models with multi-step updates. We present the \emph{Hallucinated Value Hypothesis} (HVH): updating the values of real states towards values of simulated states can result in misleading action values which adversely affect the control policy. We discuss and evaluate all four variants of Dyna amongst which three update real states toward simulated states -- so potentially toward hallucinated values -- and our proposed approach, which does not. The experimental results provide evidence for the HVH, and suggest that using predecessor models with multi-step updates is a promising direction toward developing Dyna algorithms that are more robust to model error.
format Preprint
id arxiv_https___arxiv_org_abs_2006_04363
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
Aminmansour, Farzane
Jafferjee, Taher
Imani, Ehsan
Talvitie, Erin
Bowling, Micheal
White, Martha
Machine Learning
Artificial Intelligence
Dyna-style reinforcement learning (RL) agents improve sample efficiency over model-free RL agents by updating the value function with simulated experience generated by an environment model. However, it is often difficult to learn accurate models of environment dynamics, and even small errors may result in failure of Dyna agents. In this paper, we highlight that one potential cause of that failure is bootstrapping off of the values of simulated states, and introduce a new Dyna algorithm to avoid this failure. We discuss a design space of Dyna algorithms, based on using successor or predecessor models -- simulating forwards or backwards -- and using one-step or multi-step updates. Three of the variants have been explored, but surprisingly the fourth variant has not: using predecessor models with multi-step updates. We present the \emph{Hallucinated Value Hypothesis} (HVH): updating the values of real states towards values of simulated states can result in misleading action values which adversely affect the control policy. We discuss and evaluate all four variants of Dyna amongst which three update real states toward simulated states -- so potentially toward hallucinated values -- and our proposed approach, which does not. The experimental results provide evidence for the HVH, and suggest that using predecessor models with multi-step updates is a promising direction toward developing Dyna algorithms that are more robust to model error.
title Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2006.04363