Q-value Regularized Transformer for Offline Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Shengchao, Fan, Ziqing, Huang, Chaoqin, Shen, Li, Zhang, Ya, Wang, Yanfeng, Tao, Dacheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909211296792576
author Hu, Shengchao
Fan, Ziqing
Huang, Chaoqin
Shen, Li
Zhang, Ya
Wang, Yanfeng
Tao, Dacheng
author_facet Hu, Shengchao
Fan, Ziqing
Huang, Chaoqin
Shen, Li
Zhang, Ya
Wang, Yanfeng
Tao, Dacheng
contents Recent advancements in offline reinforcement learning (RL) have underscored the capabilities of Conditional Sequence Modeling (CSM), a paradigm that learns the action distribution based on history trajectory and target returns for each state. However, these methods often struggle with stitching together optimal trajectories from sub-optimal ones due to the inconsistency between the sampled returns within individual trajectories and the optimal returns across multiple trajectories. Fortunately, Dynamic Programming (DP) methods offer a solution by leveraging a value function to approximate optimal future returns for each state, while these techniques are prone to unstable learning behaviors, particularly in long-horizon and sparse-reward scenarios. Building upon these insights, we propose the Q-value regularized Transformer (QT), which combines the trajectory modeling ability of the Transformer with the predictability of optimal future returns from DP methods. QT learns an action-value function and integrates a term maximizing action-values into the training loss of CSM, which aims to seek optimal actions that align closely with the behavior policy. Empirical evaluations on D4RL benchmark datasets demonstrate the superiority of QT over traditional DP and CSM methods, highlighting the potential of QT to enhance the state-of-the-art in offline RL.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17098
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Q-value Regularized Transformer for Offline Reinforcement Learning
Hu, Shengchao
Fan, Ziqing
Huang, Chaoqin
Shen, Li
Zhang, Ya
Wang, Yanfeng
Tao, Dacheng
Machine Learning
Recent advancements in offline reinforcement learning (RL) have underscored the capabilities of Conditional Sequence Modeling (CSM), a paradigm that learns the action distribution based on history trajectory and target returns for each state. However, these methods often struggle with stitching together optimal trajectories from sub-optimal ones due to the inconsistency between the sampled returns within individual trajectories and the optimal returns across multiple trajectories. Fortunately, Dynamic Programming (DP) methods offer a solution by leveraging a value function to approximate optimal future returns for each state, while these techniques are prone to unstable learning behaviors, particularly in long-horizon and sparse-reward scenarios. Building upon these insights, we propose the Q-value regularized Transformer (QT), which combines the trajectory modeling ability of the Transformer with the predictability of optimal future returns from DP methods. QT learns an action-value function and integrates a term maximizing action-values into the training loss of CSM, which aims to seek optimal actions that align closely with the behavior policy. Empirical evaluations on D4RL benchmark datasets demonstrate the superiority of QT over traditional DP and CSM methods, highlighting the potential of QT to enhance the state-of-the-art in offline RL.
title Q-value Regularized Transformer for Offline Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2405.17098