EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Shengjie, Liu, Shaohuai, Ye, Weirui, You, Jiacheng, Gao, Yang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914947597860864
author Wang, Shengjie
Liu, Shaohuai
Ye, Weirui
You, Jiacheng
Gao, Yang
author_facet Wang, Shengjie
Liu, Shaohuai
Ye, Weirui
You, Jiacheng
Gao, Yang
contents Sample efficiency remains a crucial challenge in applying Reinforcement Learning (RL) to real-world tasks. While recent algorithms have made significant strides in improving sample efficiency, none have achieved consistently superior performance across diverse domains. In this paper, we introduce EfficientZero V2, a general framework designed for sample-efficient RL algorithms. We have expanded the performance of EfficientZero to multiple domains, encompassing both continuous and discrete actions, as well as visual and low-dimensional inputs. With a series of improvements we propose, EfficientZero V2 outperforms the current state-of-the-art (SOTA) by a significant margin in diverse tasks under the limited data setting. EfficientZero V2 exhibits a notable advancement over the prevailing general algorithm, DreamerV3, achieving superior outcomes in 50 of 66 evaluated tasks across diverse benchmarks, such as Atari 100k, Proprio Control, and Vision Control.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00564
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data
Wang, Shengjie
Liu, Shaohuai
Ye, Weirui
You, Jiacheng
Gao, Yang
Machine Learning
Artificial Intelligence
Robotics
Sample efficiency remains a crucial challenge in applying Reinforcement Learning (RL) to real-world tasks. While recent algorithms have made significant strides in improving sample efficiency, none have achieved consistently superior performance across diverse domains. In this paper, we introduce EfficientZero V2, a general framework designed for sample-efficient RL algorithms. We have expanded the performance of EfficientZero to multiple domains, encompassing both continuous and discrete actions, as well as visual and low-dimensional inputs. With a series of improvements we propose, EfficientZero V2 outperforms the current state-of-the-art (SOTA) by a significant margin in diverse tasks under the limited data setting. EfficientZero V2 exhibits a notable advancement over the prevailing general algorithm, DreamerV3, achieving superior outcomes in 50 of 66 evaluated tasks across diverse benchmarks, such as Atari 100k, Proprio Control, and Vision Control.
title EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2403.00564