Think Before You Act: Decision Transformers with Working Memory

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kang, Jikun, Laroche, Romain, Yuan, Xingdi, Trischler, Adam, Liu, Xue, Fu, Jie
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929362776883200
author Kang, Jikun
Laroche, Romain
Yuan, Xingdi
Trischler, Adam
Liu, Xue
Fu, Jie
author_facet Kang, Jikun
Laroche, Romain
Yuan, Xingdi
Trischler, Adam
Liu, Xue
Fu, Jie
contents Decision Transformer-based decision-making agents have shown the ability to generalize across multiple tasks. However, their performance relies on massive data and computation. We argue that this inefficiency stems from the forgetting phenomenon, in which a model memorizes its behaviors in parameters throughout training. As a result, training on a new task may deteriorate the model's performance on previous tasks. In contrast to LLMs' implicit memory mechanism, the human brain utilizes distributed memory storage, which helps manage and organize multiple skills efficiently, mitigating the forgetting phenomenon. Inspired by this, we propose a working memory module to store, blend, and retrieve information for different downstream tasks. Evaluation results show that the proposed method improves training efficiency and generalization in Atari games and Meta-World object manipulation tasks. Moreover, we demonstrate that memory fine-tuning further enhances the adaptability of the proposed architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16338
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Think Before You Act: Decision Transformers with Working Memory
Kang, Jikun
Laroche, Romain
Yuan, Xingdi
Trischler, Adam
Liu, Xue
Fu, Jie
Machine Learning
Artificial Intelligence
Computation and Language
Decision Transformer-based decision-making agents have shown the ability to generalize across multiple tasks. However, their performance relies on massive data and computation. We argue that this inefficiency stems from the forgetting phenomenon, in which a model memorizes its behaviors in parameters throughout training. As a result, training on a new task may deteriorate the model's performance on previous tasks. In contrast to LLMs' implicit memory mechanism, the human brain utilizes distributed memory storage, which helps manage and organize multiple skills efficiently, mitigating the forgetting phenomenon. Inspired by this, we propose a working memory module to store, blend, and retrieve information for different downstream tasks. Evaluation results show that the proposed method improves training efficiency and generalization in Atari games and Meta-World object manipulation tasks. Moreover, we demonstrate that memory fine-tuning further enhances the adaptability of the proposed architecture.
title Think Before You Act: Decision Transformers with Working Memory
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2305.16338