Saved in:
Bibliographic Details
Main Authors: Lin, Rui, Zhang, Yiwen, Peng, Zhicheng, Lyu, Minghao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.05285
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909828722458624
author Lin, Rui
Zhang, Yiwen
Peng, Zhicheng
Lyu, Minghao
author_facet Lin, Rui
Zhang, Yiwen
Peng, Zhicheng
Lyu, Minghao
contents Decision Transformer (DT), which integrates reinforcement learning (RL) with the transformer model, introduces a novel approach to offline RL. Unlike classical algorithms that take maximizing cumulative discounted rewards as objective, DT instead maximizes the likelihood of actions. This paradigm shift, however, presents two key challenges: stitching trajectories and extrapolation of action. Existing methods, such as substituting specific tokens with predictive values and integrating the Policy Gradient (PG) method, address these challenges individually but fail to improve performance stably when combined due to inherent instability. To address this, we propose Action Gradient (AG), an innovative methodology that directly adjusts actions to fulfill a function analogous to that of PG, while also facilitating efficient integration with token prediction techniques. AG utilizes the gradient of the Q-value with respect to the action to optimize the action. The empirical results demonstrate that our method can significantly enhance the performance of DT-based algorithms, with some results achieving state-of-the-art levels.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adjusting the Output of Decision Transformer with Action Gradient
Lin, Rui
Zhang, Yiwen
Peng, Zhicheng
Lyu, Minghao
Machine Learning
Artificial Intelligence
Decision Transformer (DT), which integrates reinforcement learning (RL) with the transformer model, introduces a novel approach to offline RL. Unlike classical algorithms that take maximizing cumulative discounted rewards as objective, DT instead maximizes the likelihood of actions. This paradigm shift, however, presents two key challenges: stitching trajectories and extrapolation of action. Existing methods, such as substituting specific tokens with predictive values and integrating the Policy Gradient (PG) method, address these challenges individually but fail to improve performance stably when combined due to inherent instability. To address this, we propose Action Gradient (AG), an innovative methodology that directly adjusts actions to fulfill a function analogous to that of PG, while also facilitating efficient integration with token prediction techniques. AG utilizes the gradient of the Q-value with respect to the action to optimize the action. The empirical results demonstrate that our method can significantly enhance the performance of DT-based algorithms, with some results achieving state-of-the-art levels.
title Adjusting the Output of Decision Transformer with Action Gradient
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.05285