Hindsight Experience Replay Accelerates Proximal Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Crowder, Douglas C., McKenzie, Darrien M., Trappett, Matthew L., Chance, Frances S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Maximum Entropy Hindsight Experience Replay
by: Crowder, Douglas C., et al.
Published: (2024)
by: Crowder, Douglas C., et al.
Published: (2024)
Preliminary Tests of the Anticipatory Classifier System with Hindsight Experience Replay
by: Unold, Olgierd, et al.
Published: (2026)
by: Unold, Olgierd, et al.
Published: (2026)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
by: Romio, Gabriel, et al.
Published: (2026)
by: Romio, Gabriel, et al.
Published: (2026)
Match or Replay: Self Imitating Proximal Policy Optimization
by: Chaudhary, Gaurav, et al.
Published: (2026)
by: Chaudhary, Gaurav, et al.
Published: (2026)
Adaptable Hindsight Experience Replay for Search-Based Learning
by: Vazaios, Alexandros, et al.
Published: (2025)
by: Vazaios, Alexandros, et al.
Published: (2025)
Variance Reduction Based Experience Replay for Policy Optimization
by: Zheng, Hua, et al.
Published: (2026)
by: Zheng, Hua, et al.
Published: (2026)
Human-Aware Robot Navigation via Reinforcement Learning with Hindsight Experience Replay and Curriculum Learning
by: Li, Keyu, et al.
Published: (2021)
by: Li, Keyu, et al.
Published: (2021)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
by: Shianifar, Jonaid, et al.
Published: (2026)
by: Shianifar, Jonaid, et al.
Published: (2026)
On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
by: Zheng, Hua, et al.
Published: (2021)
by: Zheng, Hua, et al.
Published: (2021)
Layerwise Proximal Replay: A Proximal Point Method for Online Continual Learning
by: Yoo, Jason, et al.
Published: (2024)
by: Yoo, Jason, et al.
Published: (2024)
LTL-Constrained Policy Optimization with Cycle Experience Replay
by: Shah, Ameesh, et al.
Published: (2024)
by: Shah, Ameesh, et al.
Published: (2024)
Translating Flow to Policy via Hindsight Online Imitation
by: Zheng, Yitian, et al.
Published: (2025)
by: Zheng, Yitian, et al.
Published: (2025)
Revisiting Experience Replayable Conditions
by: Kobayashi, Taisuke
Published: (2024)
by: Kobayashi, Taisuke
Published: (2024)
Uncertainty Prioritized Experience Replay
by: Carrasco-Davis, Rodrigo, et al.
Published: (2025)
by: Carrasco-Davis, Rodrigo, et al.
Published: (2025)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
by: Wu, Yuning, et al.
Published: (2026)
by: Wu, Yuning, et al.
Published: (2026)
Central Path Proximal Policy Optimization
by: Milosevic, Nikola, et al.
Published: (2025)
by: Milosevic, Nikola, et al.
Published: (2025)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
by: Shankar, Kaaustaaub, et al.
Published: (2025)
by: Shankar, Kaaustaaub, et al.
Published: (2025)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025)
by: Zhong, Hai, et al.
Published: (2025)
Reliability-Adjusted Prioritized Experience Replay
by: Pleiss, Leonard S., et al.
Published: (2025)
by: Pleiss, Leonard S., et al.
Published: (2025)
RePO: Replay-Enhanced Policy Optimization
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
LatticeVision: Image to Image Networks for Modeling Non-Stationary Spatial Data
by: Sikorski, Antony, et al.
Published: (2025)
by: Sikorski, Antony, et al.
Published: (2025)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
Experience Replay with Random Reshuffling
by: Fujita, Yasuhiro
Published: (2025)
by: Fujita, Yasuhiro
Published: (2025)
Diffusion Policy through Conditional Proximal Policy Optimization
by: Liu, Ben, et al.
Published: (2026)
by: Liu, Ben, et al.
Published: (2026)
Enhancing Deep Deterministic Policy Gradients on Continuous Control Tasks with Decoupled Prioritized Experience Replay
by: Lorasdagi, Mehmet Efe, et al.
Published: (2025)
by: Lorasdagi, Mehmet Efe, et al.
Published: (2025)
Transductive Off-policy Proximal Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
Deep Gaussian Process Proximal Policy Optimization
by: van der Lende, Matthijs, et al.
Published: (2025)
by: van der Lende, Matthijs, et al.
Published: (2025)
Actor-Critic Pretraining for Proximal Policy Optimization
by: Kernbach, Andreas, et al.
Published: (2026)
by: Kernbach, Andreas, et al.
Published: (2026)
Hindsight Preference Optimization for Financial Time Series Advisory
by: Cui, Yanwei, et al.
Published: (2026)
by: Cui, Yanwei, et al.
Published: (2026)
Non-Uniform Memory Sampling in Experience Replay
by: Krutsylo, Andrii
Published: (2025)
by: Krutsylo, Andrii
Published: (2025)
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026)
by: Arnal, Charles, et al.
Published: (2026)
On the Limitation and Experience Replay for GNNs in Continual Learning
by: Su, Junwei, et al.
Published: (2023)
by: Su, Junwei, et al.
Published: (2023)
Variance Reduction via Resampling and Experience Replay
by: Han, Jiale, et al.
Published: (2025)
by: Han, Jiale, et al.
Published: (2025)
Beyond the Boundaries of Proximal Policy Optimization
by: Tan, Charlie B., et al.
Published: (2024)
by: Tan, Charlie B., et al.
Published: (2024)
Proximal Policy Optimization with Adaptive Exploration
by: Lixandru, Andrei
Published: (2024)
by: Lixandru, Andrei
Published: (2024)
Complexity-Regularized Proximal Policy Optimization
by: Serfilippi, Luca, et al.
Published: (2025)
by: Serfilippi, Luca, et al.
Published: (2025)
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
by: Ahmad, Ahmad, et al.
Published: (2024)
by: Ahmad, Ahmad, et al.
Published: (2024)
Token-level Proximal Policy Optimization for Query Generation
by: Ouyang, Yichen, et al.
Published: (2024)
by: Ouyang, Yichen, et al.
Published: (2024)
Similar Items
-
Maximum Entropy Hindsight Experience Replay
by: Crowder, Douglas C., et al.
Published: (2024) -
Preliminary Tests of the Anticipatory Classifier System with Hindsight Experience Replay
by: Unold, Olgierd, et al.
Published: (2026) -
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
by: Romio, Gabriel, et al.
Published: (2026) -
Match or Replay: Self Imitating Proximal Policy Optimization
by: Chaudhary, Gaurav, et al.
Published: (2026) -
Adaptable Hindsight Experience Replay for Search-Based Learning
by: Vazaios, Alexandros, et al.
Published: (2025)