Salvato in:
Dettagli Bibliografici
Autori principali: Lei, Xing, Yang, Wenyan, Ke, Kaiqiang, Yang, Shentao, Zhang, Xuetao, Pajarinen, Joni, Wang, Donglin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2508.06108
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918119109296128
author Lei, Xing
Yang, Wenyan
Ke, Kaiqiang
Yang, Shentao
Zhang, Xuetao
Pajarinen, Joni
Wang, Donglin
author_facet Lei, Xing
Yang, Wenyan
Ke, Kaiqiang
Yang, Shentao
Zhang, Xuetao
Pajarinen, Joni
Wang, Donglin
contents Goal-conditioned reinforcement learning (GCRL) with sparse rewards remains a fundamental challenge in reinforcement learning. While hindsight experience replay (HER) has shown promise by relabeling collected trajectories with achieved goals, we argue that trajectory relabeling alone does not fully exploit the available experiences in off-policy GCRL methods, resulting in limited sample efficiency. In this paper, we propose Hindsight Goal-conditioned Regularization (HGR), a technique that generates action regularization priors based on hindsight goals. When combined with hindsight self-imitation regularization (HSR), our approach enables off-policy RL algorithms to maximize experience utilization. Compared to existing GCRL methods that employ HER and self-imitation techniques, our hindsight regularizations achieve substantially more efficient sample reuse and the best performances, which we empirically demonstrate on a suite of navigation and manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning
Lei, Xing
Yang, Wenyan
Ke, Kaiqiang
Yang, Shentao
Zhang, Xuetao
Pajarinen, Joni
Wang, Donglin
Machine Learning
Artificial Intelligence
Goal-conditioned reinforcement learning (GCRL) with sparse rewards remains a fundamental challenge in reinforcement learning. While hindsight experience replay (HER) has shown promise by relabeling collected trajectories with achieved goals, we argue that trajectory relabeling alone does not fully exploit the available experiences in off-policy GCRL methods, resulting in limited sample efficiency. In this paper, we propose Hindsight Goal-conditioned Regularization (HGR), a technique that generates action regularization priors based on hindsight goals. When combined with hindsight self-imitation regularization (HSR), our approach enables off-policy RL algorithms to maximize experience utilization. Compared to existing GCRL methods that employ HER and self-imitation techniques, our hindsight regularizations achieve substantially more efficient sample reuse and the best performances, which we empirically demonstrate on a suite of navigation and manipulation tasks.
title GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.06108