Hindsight Experience Replay Accelerates Proximal Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Crowder, Douglas C., McKenzie, Darrien M., Trappett, Matthew L., Chance, Frances S.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567323717632
author Crowder, Douglas C.
McKenzie, Darrien M.
Trappett, Matthew L.
Chance, Frances S.
author_facet Crowder, Douglas C.
McKenzie, Darrien M.
Trappett, Matthew L.
Chance, Frances S.
contents Hindsight experience replay (HER) accelerates off-policy reinforcement learning algorithms for environments that emit sparse rewards by modifying the goal of the episode post-hoc to be some state achieved during the episode. Because post-hoc modification of the observed goal violates the assumptions of on-policy algorithms, HER is not typically applied to on-policy algorithms. Here, we show that HER can dramatically accelerate proximal policy optimization (PPO), an on-policy reinforcement learning algorithm, when tested on a custom predator-prey environment.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22524
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hindsight Experience Replay Accelerates Proximal Policy Optimization
Crowder, Douglas C.
McKenzie, Darrien M.
Trappett, Matthew L.
Chance, Frances S.
Machine Learning
Hindsight experience replay (HER) accelerates off-policy reinforcement learning algorithms for environments that emit sparse rewards by modifying the goal of the episode post-hoc to be some state achieved during the episode. Because post-hoc modification of the observed goal violates the assumptions of on-policy algorithms, HER is not typically applied to on-policy algorithms. Here, we show that HER can dramatically accelerate proximal policy optimization (PPO), an on-policy reinforcement learning algorithm, when tested on a custom predator-prey environment.
title Hindsight Experience Replay Accelerates Proximal Policy Optimization
topic Machine Learning
url https://arxiv.org/abs/2410.22524