In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Sili, Hu, Jifeng, Chen, Hechang, Sun, Lichao, Yang, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916267431034880
author Huang, Sili
Hu, Jifeng
Chen, Hechang
Sun, Lichao
Yang, Bo
author_facet Huang, Sili
Hu, Jifeng
Chen, Hechang
Sun, Lichao
Yang, Bo
contents In-context learning is a promising approach for offline reinforcement learning (RL) to handle online tasks, which can be achieved by providing task prompts. Recent works demonstrated that in-context RL could emerge with self-improvement in a trial-and-error manner when treating RL tasks as an across-episodic sequential prediction problem. Despite the self-improvement not requiring gradient updates, current works still suffer from high computational costs when the across-episodic sequence increases with task horizons. To this end, we propose an In-context Decision Transformer (IDT) to achieve self-improvement in a high-level trial-and-error manner. Specifically, IDT is inspired by the efficient hierarchical structure of human decision-making and thus reconstructs the sequence to consist of high-level decisions instead of low-level actions that interact with environments. As one high-level decision can guide multi-step low-level actions, IDT naturally avoids excessively long sequences and solves online tasks more efficiently. Experimental results show that IDT achieves state-of-the-art in long-horizon tasks over current in-context RL methods. In particular, the online evaluation time of our IDT is \textbf{36$\times$} times faster than baselines in the D4RL benchmark and \textbf{27$\times$} times faster in the Grid World benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20692
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
Huang, Sili
Hu, Jifeng
Chen, Hechang
Sun, Lichao
Yang, Bo
Machine Learning
Artificial Intelligence
In-context learning is a promising approach for offline reinforcement learning (RL) to handle online tasks, which can be achieved by providing task prompts. Recent works demonstrated that in-context RL could emerge with self-improvement in a trial-and-error manner when treating RL tasks as an across-episodic sequential prediction problem. Despite the self-improvement not requiring gradient updates, current works still suffer from high computational costs when the across-episodic sequence increases with task horizons. To this end, we propose an In-context Decision Transformer (IDT) to achieve self-improvement in a high-level trial-and-error manner. Specifically, IDT is inspired by the efficient hierarchical structure of human decision-making and thus reconstructs the sequence to consist of high-level decisions instead of low-level actions that interact with environments. As one high-level decision can guide multi-step low-level actions, IDT naturally avoids excessively long sequences and solves online tasks more efficiently. Experimental results show that IDT achieves state-of-the-art in long-horizon tasks over current in-context RL methods. In particular, the online evaluation time of our IDT is \textbf{36$\times$} times faster than baselines in the D4RL benchmark and \textbf{27$\times$} times faster in the Grid World benchmark.
title In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.20692