Saved in:
Bibliographic Details
Main Authors: Barkley, Brett, Fridovich-Keil, David
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.14312
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918066543132672
author Barkley, Brett
Fridovich-Keil, David
author_facet Barkley, Brett
Fridovich-Keil, David
contents Dyna-style off-policy model-based reinforcement learning (DMBRL) algorithms are a family of techniques for generating synthetic state transition data and thereby enhancing the sample efficiency of off-policy RL algorithms. This paper identifies and investigates a surprising performance gap observed when applying DMBRL algorithms across different benchmark environments with proprioceptive observations. We show that, while DMBRL algorithms perform well in OpenAI Gym, their performance can drop significantly in DeepMind Control Suite (DMC), even though these settings offer similar tasks and identical physics backends. Modern techniques designed to address several key issues that arise in these settings do not provide a consistent improvement across all environments, and overall our results show that adding synthetic rollouts to the training process -- the backbone of Dyna-style algorithms -- significantly degrades performance across most DMC environments. Our findings contribute to a deeper understanding of several fundamental challenges in model-based RL and show that, like many optimization fields, there is no free lunch when evaluating performance across diverse benchmarks in RL.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14312
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning
Barkley, Brett
Fridovich-Keil, David
Machine Learning
Dyna-style off-policy model-based reinforcement learning (DMBRL) algorithms are a family of techniques for generating synthetic state transition data and thereby enhancing the sample efficiency of off-policy RL algorithms. This paper identifies and investigates a surprising performance gap observed when applying DMBRL algorithms across different benchmark environments with proprioceptive observations. We show that, while DMBRL algorithms perform well in OpenAI Gym, their performance can drop significantly in DeepMind Control Suite (DMC), even though these settings offer similar tasks and identical physics backends. Modern techniques designed to address several key issues that arise in these settings do not provide a consistent improvement across all environments, and overall our results show that adding synthetic rollouts to the training process -- the backbone of Dyna-style algorithms -- significantly degrades performance across most DMC environments. Our findings contribute to a deeper understanding of several fundamental challenges in model-based RL and show that, like many optimization fields, there is no free lunch when evaluating performance across diverse benchmarks in RL.
title Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2412.14312