When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yansong, Dong, Zeyu, Luo, Ertai, Wu, Yu, Wu, Shuo, Han, Shuo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909351559561216
author Li, Yansong
Dong, Zeyu
Luo, Ertai
Wu, Yu
Wu, Shuo
Han, Shuo
author_facet Li, Yansong
Dong, Zeyu
Luo, Ertai
Wu, Yu
Wu, Shuo
Han, Shuo
contents Reinforcement learning (RL) algorithms can be divided into two classes: model-free algorithms, which are sample-inefficient, and model-based algorithms, which suffer from model bias. Dyna-style algorithms combine these two approaches by using simulated data from an estimated environmental model to accelerate model-free training. However, their efficiency is compromised when the estimated model is inaccurate. Previous works address this issue by using model ensembles or pretraining the estimated model with data collected from the real environment, increasing computational and sample complexity. To tackle this issue, we introduce an out-of-distribution (OOD) data filter that removes simulated data from the estimated model that significantly diverges from data collected in the real environment. We show theoretically that this technique enhances the quality of simulated data. With the help of the OOD data filter, the data simulated from the estimated model better mimics the data collected by interacting with the real model. This improvement is evident in the critic updates compared to using the simulated data without the OOD data filter. Our experiment integrates the data filter into the model-based policy optimization (MBPO) algorithm. The results demonstrate that our method requires fewer interactions with the real environment to achieve a higher level of optimality than MBPO, even without a model ensemble.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter
Li, Yansong
Dong, Zeyu
Luo, Ertai
Wu, Yu
Wu, Shuo
Han, Shuo
Machine Learning
Systems and Control
Reinforcement learning (RL) algorithms can be divided into two classes: model-free algorithms, which are sample-inefficient, and model-based algorithms, which suffer from model bias. Dyna-style algorithms combine these two approaches by using simulated data from an estimated environmental model to accelerate model-free training. However, their efficiency is compromised when the estimated model is inaccurate. Previous works address this issue by using model ensembles or pretraining the estimated model with data collected from the real environment, increasing computational and sample complexity. To tackle this issue, we introduce an out-of-distribution (OOD) data filter that removes simulated data from the estimated model that significantly diverges from data collected in the real environment. We show theoretically that this technique enhances the quality of simulated data. With the help of the OOD data filter, the data simulated from the estimated model better mimics the data collected by interacting with the real model. This improvement is evident in the critic updates compared to using the simulated data without the OOD data filter. Our experiment integrates the data filter into the model-based policy optimization (MBPO) algorithm. The results demonstrate that our method requires fewer interactions with the real environment to achieve a higher level of optimality than MBPO, even without a model ensemble.
title When to Trust Your Data: Enhancing Dyna-Style Model-Based Reinforcement Learning With Data Filter
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2410.12160