Enhancing Decision Transformer with Diffusion-Based Trajectory Branch Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhihong, Qian, Long, Liu, Zeyang, Wan, Lipeng, Chen, Xingyu, Lan, Xuguang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913579934941184
author Liu, Zhihong
Qian, Long
Liu, Zeyang
Wan, Lipeng
Chen, Xingyu
Lan, Xuguang
author_facet Liu, Zhihong
Qian, Long
Liu, Zeyang
Wan, Lipeng
Chen, Xingyu
Lan, Xuguang
contents Decision Transformer (DT) can learn effective policy from offline datasets by converting the offline reinforcement learning (RL) into a supervised sequence modeling task, where the trajectory elements are generated auto-regressively conditioned on the return-to-go (RTG).However, the sequence modeling learning approach tends to learn policies that converge on the sub-optimal trajectories within the dataset, for lack of bridging data to move to better trajectories, even if the condition is set to the highest RTG.To address this issue, we introduce Diffusion-Based Trajectory Branch Generation (BG), which expands the trajectories of the dataset with branches generated by a diffusion model.The trajectory branch is generated based on the segment of the trajectory within the dataset, and leads to trajectories with higher returns.We concatenate the generated branch with the trajectory segment as an expansion of the trajectory.After expanding, DT has more opportunities to learn policies to move to better trajectories, preventing it from converging to the sub-optimal trajectories.Empirically, after processing with BG, DT outperforms state-of-the-art sequence modeling methods on D4RL benchmark, demonstrating the effectiveness of adding branches to the dataset without further modifications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11327
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Decision Transformer with Diffusion-Based Trajectory Branch Generation
Liu, Zhihong
Qian, Long
Liu, Zeyang
Wan, Lipeng
Chen, Xingyu
Lan, Xuguang
Machine Learning
Decision Transformer (DT) can learn effective policy from offline datasets by converting the offline reinforcement learning (RL) into a supervised sequence modeling task, where the trajectory elements are generated auto-regressively conditioned on the return-to-go (RTG).However, the sequence modeling learning approach tends to learn policies that converge on the sub-optimal trajectories within the dataset, for lack of bridging data to move to better trajectories, even if the condition is set to the highest RTG.To address this issue, we introduce Diffusion-Based Trajectory Branch Generation (BG), which expands the trajectories of the dataset with branches generated by a diffusion model.The trajectory branch is generated based on the segment of the trajectory within the dataset, and leads to trajectories with higher returns.We concatenate the generated branch with the trajectory segment as an expansion of the trajectory.After expanding, DT has more opportunities to learn policies to move to better trajectories, preventing it from converging to the sub-optimal trajectories.Empirically, after processing with BG, DT outperforms state-of-the-art sequence modeling methods on D4RL benchmark, demonstrating the effectiveness of adding branches to the dataset without further modifications.
title Enhancing Decision Transformer with Diffusion-Based Trajectory Branch Generation
topic Machine Learning
url https://arxiv.org/abs/2411.11327