Decoupling Return-to-Go for Efficient Decision Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yongyi, Liu, Hanyu, Li, Lingfeng, Chen, Bozhou, Li, Ang, Zheng, Qirui, Yang, Xionghui, Li, Wenxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909997911244800
author Wang, Yongyi
Liu, Hanyu
Li, Lingfeng
Chen, Bozhou
Li, Ang
Zheng, Qirui
Yang, Xionghui
Li, Wenxin
author_facet Wang, Yongyi
Liu, Hanyu
Li, Lingfeng
Chen, Bozhou
Li, Ang
Zheng, Qirui
Yang, Xionghui
Li, Wenxin
contents The Decision Transformer (DT) has established a powerful sequence modeling approach to offline reinforcement learning. It conditions its action predictions on Return-to-Go (RTG), using it both to distinguish trajectory quality during training and to guide action generation at inference. In this work, we identify a critical redundancy in this design: feeding the entire sequence of RTGs into the Transformer is theoretically unnecessary, as only the most recent RTG affects action prediction. We show that this redundancy can impair DT's performance through experiments. To resolve this, we propose the Decoupled DT (DDT). DDT simplifies the architecture by processing only observation and action sequences through the Transformer, using the latest RTG to guide the action prediction. This streamlined approach not only improves performance but also reduces computational cost. Our experiments show that DDT significantly outperforms DT and establishes competitive performance against state-of-the-art DT variants across multiple offline RL tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15953
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Decoupling Return-to-Go for Efficient Decision Transformer
Wang, Yongyi
Liu, Hanyu
Li, Lingfeng
Chen, Bozhou
Li, Ang
Zheng, Qirui
Yang, Xionghui
Li, Wenxin
Artificial Intelligence
The Decision Transformer (DT) has established a powerful sequence modeling approach to offline reinforcement learning. It conditions its action predictions on Return-to-Go (RTG), using it both to distinguish trajectory quality during training and to guide action generation at inference. In this work, we identify a critical redundancy in this design: feeding the entire sequence of RTGs into the Transformer is theoretically unnecessary, as only the most recent RTG affects action prediction. We show that this redundancy can impair DT's performance through experiments. To resolve this, we propose the Decoupled DT (DDT). DDT simplifies the architecture by processing only observation and action sequences through the Transformer, using the latest RTG to guide the action prediction. This streamlined approach not only improves performance but also reduces computational cost. Our experiments show that DDT significantly outperforms DT and establishes competitive performance against state-of-the-art DT variants across multiple offline RL tasks.
title Decoupling Return-to-Go for Efficient Decision Transformer
topic Artificial Intelligence
url https://arxiv.org/abs/2601.15953