Hindsight Experience Replay Accelerates Proximal Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913567323717632 |
|---|---|
| author | Crowder, Douglas C. McKenzie, Darrien M. Trappett, Matthew L. Chance, Frances S. |
| author_facet | Crowder, Douglas C. McKenzie, Darrien M. Trappett, Matthew L. Chance, Frances S. |
| contents | Hindsight experience replay (HER) accelerates off-policy reinforcement learning algorithms for environments that emit sparse rewards by modifying the goal of the episode post-hoc to be some state achieved during the episode. Because post-hoc modification of the observed goal violates the assumptions of on-policy algorithms, HER is not typically applied to on-policy algorithms. Here, we show that HER can dramatically accelerate proximal policy optimization (PPO), an on-policy reinforcement learning algorithm, when tested on a custom predator-prey environment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_22524 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Hindsight Experience Replay Accelerates Proximal Policy Optimization Crowder, Douglas C. McKenzie, Darrien M. Trappett, Matthew L. Chance, Frances S. Machine Learning Hindsight experience replay (HER) accelerates off-policy reinforcement learning algorithms for environments that emit sparse rewards by modifying the goal of the episode post-hoc to be some state achieved during the episode. Because post-hoc modification of the observed goal violates the assumptions of on-policy algorithms, HER is not typically applied to on-policy algorithms. Here, we show that HER can dramatically accelerate proximal policy optimization (PPO), an on-policy reinforcement learning algorithm, when tested on a custom predator-prey environment. |
| title | Hindsight Experience Replay Accelerates Proximal Policy Optimization |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2410.22524 |