Decoupling Return-to-Go for Efficient Decision Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909997911244800 |
|---|---|
| author | Wang, Yongyi Liu, Hanyu Li, Lingfeng Chen, Bozhou Li, Ang Zheng, Qirui Yang, Xionghui Li, Wenxin |
| author_facet | Wang, Yongyi Liu, Hanyu Li, Lingfeng Chen, Bozhou Li, Ang Zheng, Qirui Yang, Xionghui Li, Wenxin |
| contents | The Decision Transformer (DT) has established a powerful sequence modeling approach to offline reinforcement learning. It conditions its action predictions on Return-to-Go (RTG), using it both to distinguish trajectory quality during training and to guide action generation at inference. In this work, we identify a critical redundancy in this design: feeding the entire sequence of RTGs into the Transformer is theoretically unnecessary, as only the most recent RTG affects action prediction. We show that this redundancy can impair DT's performance through experiments. To resolve this, we propose the Decoupled DT (DDT). DDT simplifies the architecture by processing only observation and action sequences through the Transformer, using the latest RTG to guide the action prediction. This streamlined approach not only improves performance but also reduces computational cost. Our experiments show that DDT significantly outperforms DT and establishes competitive performance against state-of-the-art DT variants across multiple offline RL tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_15953 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Decoupling Return-to-Go for Efficient Decision Transformer Wang, Yongyi Liu, Hanyu Li, Lingfeng Chen, Bozhou Li, Ang Zheng, Qirui Yang, Xionghui Li, Wenxin Artificial Intelligence The Decision Transformer (DT) has established a powerful sequence modeling approach to offline reinforcement learning. It conditions its action predictions on Return-to-Go (RTG), using it both to distinguish trajectory quality during training and to guide action generation at inference. In this work, we identify a critical redundancy in this design: feeding the entire sequence of RTGs into the Transformer is theoretically unnecessary, as only the most recent RTG affects action prediction. We show that this redundancy can impair DT's performance through experiments. To resolve this, we propose the Decoupled DT (DDT). DDT simplifies the architecture by processing only observation and action sequences through the Transformer, using the latest RTG to guide the action prediction. This streamlined approach not only improves performance but also reduces computational cost. Our experiments show that DDT significantly outperforms DT and establishes competitive performance against state-of-the-art DT variants across multiple offline RL tasks. |
| title | Decoupling Return-to-Go for Efficient Decision Transformer |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2601.15953 |