Open-World Reinforcement Learning over Long Short-Term Imagination

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Jiajian, Wang, Qi, Wang, Yunbo, Jin, Xin, Li, Yang, Zeng, Wenjun, Yang, Xiaokang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911494154747904
author Li, Jiajian
Wang, Qi
Wang, Yunbo
Jin, Xin
Li, Yang
Zeng, Wenjun
Yang, Xiaokang
author_facet Li, Jiajian
Wang, Qi
Wang, Yunbo
Jin, Xin
Li, Yang
Zeng, Wenjun
Yang, Xiaokang
contents Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be "short-sighted", as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03618
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open-World Reinforcement Learning over Long Short-Term Imagination
Li, Jiajian
Wang, Qi
Wang, Yunbo
Jin, Xin
Li, Yang
Zeng, Wenjun
Yang, Xiaokang
Machine Learning
Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be "short-sighted", as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo.
title Open-World Reinforcement Learning over Long Short-Term Imagination
topic Machine Learning
url https://arxiv.org/abs/2410.03618