Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cheng, Jie, Qiao, Ruixi, Ma, Yingwei, Li, Binhua, Xiong, Gang, Miao, Qinghai, Li, Yongbin, Lv, Yisheng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911406372159488
author Cheng, Jie
Qiao, Ruixi
Ma, Yingwei
Li, Binhua
Xiong, Gang
Miao, Qinghai
Li, Yongbin
Lv, Yisheng
author_facet Cheng, Jie
Qiao, Ruixi
Ma, Yingwei
Li, Binhua
Xiong, Gang
Miao, Qinghai
Li, Yongbin
Lv, Yisheng
contents A significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks. Inspired by the excellent generalization of world model in conditional video generation, we explore the potential of image observation-based world model for scaling offline RL and enhancing generalization on novel tasks. In this paper, we introduce JOWA: Jointly-Optimized World-Action model, an offline model-based RL agent pretrained on multiple Atari games with 6 billion tokens data to learn general-purpose representation and decision-making ability. Our method jointly optimizes a world-action model through a shared transformer backbone, which stabilize temporal difference learning with large models during pretraining. Moreover, we propose a provably efficient and parallelizable planning algorithm to compensate for the Q-value estimation error and thus search out better policies. Experimental results indicate that our largest agent, with 150 million parameters, achieves 78.9% human-level performance on pretrained games using only 10% subsampled offline data, outperforming existing state-of-the-art large-scale offline RL baselines by 31.6% on averange. Furthermore, JOWA scales favorably with model capacity and can sample-efficiently transfer to novel games using only 5k offline fine-tuning data (approximately 4 trajectories) per game, demonstrating superior generalization. We will release codes and model weights at https://github.com/CJReinforce/JOWA
format Preprint
id arxiv_https___arxiv_org_abs_2410_00564
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
Cheng, Jie
Qiao, Ruixi
Ma, Yingwei
Li, Binhua
Xiong, Gang
Miao, Qinghai
Li, Yongbin
Lv, Yisheng
Machine Learning
Artificial Intelligence
A significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks. Inspired by the excellent generalization of world model in conditional video generation, we explore the potential of image observation-based world model for scaling offline RL and enhancing generalization on novel tasks. In this paper, we introduce JOWA: Jointly-Optimized World-Action model, an offline model-based RL agent pretrained on multiple Atari games with 6 billion tokens data to learn general-purpose representation and decision-making ability. Our method jointly optimizes a world-action model through a shared transformer backbone, which stabilize temporal difference learning with large models during pretraining. Moreover, we propose a provably efficient and parallelizable planning algorithm to compensate for the Q-value estimation error and thus search out better policies. Experimental results indicate that our largest agent, with 150 million parameters, achieves 78.9% human-level performance on pretrained games using only 10% subsampled offline data, outperforming existing state-of-the-art large-scale offline RL baselines by 31.6% on averange. Furthermore, JOWA scales favorably with model capacity and can sample-efficiently transfer to novel games using only 5k offline fine-tuning data (approximately 4 trajectories) per game, demonstrating superior generalization. We will release codes and model weights at https://github.com/CJReinforce/JOWA
title Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.00564