ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Chen-Xiao, Wu, Chenyang, Cao, Mingjun, Kong, Rui, Zhang, Zongzhang, Yu, Yang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914845777985536
author Gao, Chen-Xiao
Wu, Chenyang
Cao, Mingjun
Kong, Rui
Zhang, Zongzhang
Yu, Yang
author_facet Gao, Chen-Xiao
Wu, Chenyang
Cao, Mingjun
Kong, Rui
Zhang, Zongzhang
Yu, Yang
contents Decision Transformer (DT), which employs expressive sequence modeling techniques to perform action generation, has emerged as a promising approach to offline policy optimization. However, DT generates actions conditioned on a desired future return, which is known to bear some weaknesses such as the susceptibility to environmental stochasticity. To overcome DT's weaknesses, we propose to empower DT with dynamic programming. Our method comprises three steps. First, we employ in-sample value iteration to obtain approximated value functions, which involves dynamic programming over the MDP structure. Second, we evaluate action quality in context with estimated advantages. We introduce two types of advantage estimators, IAE and GAE, which are suitable for different tasks. Third, we train an Advantage-Conditioned Transformer (ACT) to generate actions conditioned on the estimated advantages. Finally, during testing, ACT generates actions conditioned on a desired advantage. Our evaluation results validate that, by leveraging the power of dynamic programming, ACT demonstrates effective trajectory stitching and robust action generation in spite of the environmental stochasticity, outperforming baseline methods across various benchmarks. Additionally, we conduct an in-depth analysis of ACT's various design choices through ablation studies. Our code is available at https://github.com/LAMDA-RL/ACT.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05915
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning
Gao, Chen-Xiao
Wu, Chenyang
Cao, Mingjun
Kong, Rui
Zhang, Zongzhang
Yu, Yang
Machine Learning
Artificial Intelligence
Decision Transformer (DT), which employs expressive sequence modeling techniques to perform action generation, has emerged as a promising approach to offline policy optimization. However, DT generates actions conditioned on a desired future return, which is known to bear some weaknesses such as the susceptibility to environmental stochasticity. To overcome DT's weaknesses, we propose to empower DT with dynamic programming. Our method comprises three steps. First, we employ in-sample value iteration to obtain approximated value functions, which involves dynamic programming over the MDP structure. Second, we evaluate action quality in context with estimated advantages. We introduce two types of advantage estimators, IAE and GAE, which are suitable for different tasks. Third, we train an Advantage-Conditioned Transformer (ACT) to generate actions conditioned on the estimated advantages. Finally, during testing, ACT generates actions conditioned on a desired advantage. Our evaluation results validate that, by leveraging the power of dynamic programming, ACT demonstrates effective trajectory stitching and robust action generation in spite of the environmental stochasticity, outperforming baseline methods across various benchmarks. Additionally, we conduct an in-depth analysis of ACT's various design choices through ablation studies. Our code is available at https://github.com/LAMDA-RL/ACT.
title ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2309.05915