Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Seo, Younggyo, Abbeel, Pieter
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912711201259520
author Seo, Younggyo
Abbeel, Pieter
author_facet Seo, Younggyo
Abbeel, Pieter
contents Predicting a sequence of actions has been crucial in the success of recent behavior cloning algorithms in robotics. Can similar ideas improve reinforcement learning (RL)? We answer affirmatively by observing that incorporating action sequences when predicting ground-truth return-to-go leads to lower validation loss. Motivated by this, we introduce Coarse-to-fine Q-Network with Action Sequence (CQN-AS), a novel value-based RL algorithm that learns a critic network that outputs Q-values over a sequence of actions, i.e., explicitly training the value function to learn the consequence of executing action sequences. Our experiments show that CQN-AS outperforms several baselines on a variety of sparse-reward humanoid control and tabletop manipulation tasks from BiGym and RLBench.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
Seo, Younggyo
Abbeel, Pieter
Machine Learning
Artificial Intelligence
Robotics
Predicting a sequence of actions has been crucial in the success of recent behavior cloning algorithms in robotics. Can similar ideas improve reinforcement learning (RL)? We answer affirmatively by observing that incorporating action sequences when predicting ground-truth return-to-go leads to lower validation loss. Motivated by this, we introduce Coarse-to-fine Q-Network with Action Sequence (CQN-AS), a novel value-based RL algorithm that learns a critic network that outputs Q-values over a sequence of actions, i.e., explicitly training the value function to learn the consequence of executing action sequences. Our experiments show that CQN-AS outperforms several baselines on a variety of sparse-reward humanoid control and tabletop manipulation tasks from BiGym and RLBench.
title Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2411.12155