Saved in:
Bibliographic Details
Main Authors: Wu, Qingyuan, Zhan, Simon Sinong, Wang, Yixuan, Wang, Yuhui, Lin, Chung-Wei, Lv, Chen, Zhu, Qi, Schmidhuber, Jürgen, Huang, Chao
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.03141
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916276433059840
author Wu, Qingyuan
Zhan, Simon Sinong
Wang, Yixuan
Wang, Yuhui
Lin, Chung-Wei
Lv, Chen
Zhu, Qi
Schmidhuber, Jürgen
Huang, Chao
author_facet Wu, Qingyuan
Zhan, Simon Sinong
Wang, Yixuan
Wang, Yuhui
Lin, Chung-Wei
Lv, Chen
Zhu, Qi
Schmidhuber, Jürgen
Huang, Chao
contents Reinforcement learning (RL) is challenging in the common case of delays between events and their sensory perceptions. State-of-the-art (SOTA) state augmentation techniques either suffer from state space explosion or performance degeneration in stochastic environments. To address these challenges, we present a novel Auxiliary-Delayed Reinforcement Learning (AD-RL) method that leverages auxiliary tasks involving short delays to accelerate RL with long delays, without compromising performance in stochastic environments. Specifically, AD-RL learns a value function for short delays and uses bootstrapping and policy improvement techniques to adjust it for long delays. We theoretically show that this can greatly reduce the sample complexity. On deterministic and stochastic benchmarks, our method significantly outperforms the SOTAs in both sample efficiency and policy performance. Code is available at https://github.com/QingyuanWuNothing/AD-RL.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03141
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays
Wu, Qingyuan
Zhan, Simon Sinong
Wang, Yixuan
Wang, Yuhui
Lin, Chung-Wei
Lv, Chen
Zhu, Qi
Schmidhuber, Jürgen
Huang, Chao
Machine Learning
Artificial Intelligence
Systems and Control
Reinforcement learning (RL) is challenging in the common case of delays between events and their sensory perceptions. State-of-the-art (SOTA) state augmentation techniques either suffer from state space explosion or performance degeneration in stochastic environments. To address these challenges, we present a novel Auxiliary-Delayed Reinforcement Learning (AD-RL) method that leverages auxiliary tasks involving short delays to accelerate RL with long delays, without compromising performance in stochastic environments. Specifically, AD-RL learns a value function for short delays and uses bootstrapping and policy improvement techniques to adjust it for long delays. We theoretically show that this can greatly reduce the sample complexity. On deterministic and stochastic benchmarks, our method significantly outperforms the SOTAs in both sample efficiency and policy performance. Code is available at https://github.com/QingyuanWuNothing/AD-RL.
title Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays
topic Machine Learning
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2402.03141