Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Guwen, Sato, Masahiro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
by: Romio, Gabriel, et al.
Published: (2026)
by: Romio, Gabriel, et al.
Published: (2026)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
by: Cho, Minjae, et al.
Published: (2026)
by: Cho, Minjae, et al.
Published: (2026)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
Experience Replay with Random Reshuffling
by: Fujita, Yasuhiro
Published: (2025)
by: Fujita, Yasuhiro
Published: (2025)
Sample Efficient Experience Replay in Non-stationary Environments
by: Duan, Tianyang, et al.
Published: (2025)
by: Duan, Tianyang, et al.
Published: (2025)
ROER: Regularized Optimal Experience Replay
by: Li, Changling, et al.
Published: (2024)
by: Li, Changling, et al.
Published: (2024)
R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
by: Jiang, Zhizheng, et al.
Published: (2026)
by: Jiang, Zhizheng, et al.
Published: (2026)
Mastering the Game of Go with Self-play Experience Replay
by: Liu, Jingbin, et al.
Published: (2026)
by: Liu, Jingbin, et al.
Published: (2026)
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
by: Chen, Yuyang, et al.
Published: (2024)
by: Chen, Yuyang, et al.
Published: (2024)
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
by: Zhao, Kaiyan, et al.
Published: (2024)
by: Zhao, Kaiyan, et al.
Published: (2024)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
by: Holmes, Ian, et al.
Published: (2025)
by: Holmes, Ian, et al.
Published: (2025)
Investigating the Interplay of Prioritized Replay and Generalization
by: Panahi, Parham Mohammad, et al.
Published: (2024)
by: Panahi, Parham Mohammad, et al.
Published: (2024)
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
by: Cameron, Chris, et al.
Published: (2026)
by: Cameron, Chris, et al.
Published: (2026)
DyJR: Preserving Diversity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
by: Li, Long, et al.
Published: (2026)
by: Li, Long, et al.
Published: (2026)
Better Generative Replay for Continual Federated Learning
by: Qi, Daiqing, et al.
Published: (2023)
by: Qi, Daiqing, et al.
Published: (2023)
CIER: A Novel Experience Replay Approach with Causal Inference in Deep Reinforcement Learning
by: Wang, Jingwen, et al.
Published: (2024)
by: Wang, Jingwen, et al.
Published: (2024)
On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
by: Zheng, Hua, et al.
Published: (2021)
by: Zheng, Hua, et al.
Published: (2021)
What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning?
by: Shihab, Ibne Farabi, et al.
Published: (2025)
by: Shihab, Ibne Farabi, et al.
Published: (2025)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
by: Rezaei, Mohammad, et al.
Published: (2026)
by: Rezaei, Mohammad, et al.
Published: (2026)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
by: Yenicesu, Arda Sarp, et al.
Published: (2024)
by: Yenicesu, Arda Sarp, et al.
Published: (2024)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Multi-Turn Code Generation Through Single-Step Rewards
by: Jain, Arnav Kumar, et al.
Published: (2025)
by: Jain, Arnav Kumar, et al.
Published: (2025)
Experience Replay Addresses Loss of Plasticity in Continual Learning
by: Wang, Jiuqi, et al.
Published: (2025)
by: Wang, Jiuqi, et al.
Published: (2025)
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
by: Lei, Shiye, et al.
Published: (2026)
by: Lei, Shiye, et al.
Published: (2026)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
by: Liu, Jinyi, et al.
Published: (2023)
by: Liu, Jinyi, et al.
Published: (2023)
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning
by: Ma, Yiran, et al.
Published: (2024)
by: Ma, Yiran, et al.
Published: (2024)
MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels
by: Luo, Tianyang, et al.
Published: (2026)
by: Luo, Tianyang, et al.
Published: (2026)
Data-Free Generative Replay for Class-Incremental Learning on Imbalanced Data
by: Younis, Sohaib, et al.
Published: (2024)
by: Younis, Sohaib, et al.
Published: (2024)
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
by: Yoon, Deokgyu, et al.
Published: (2026)
by: Yoon, Deokgyu, et al.
Published: (2026)
Human-Aware Robot Navigation via Reinforcement Learning with Hindsight Experience Replay and Curriculum Learning
by: Li, Keyu, et al.
Published: (2021)
by: Li, Keyu, et al.
Published: (2021)
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026)
by: Liu, Zhanwang, et al.
Published: (2026)
2048: Reinforcement Learning in a Delayed Reward Environment
by: Saligram, Prady, et al.
Published: (2025)
by: Saligram, Prady, et al.
Published: (2025)
Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay
by: Liu, Jinmei, et al.
Published: (2024)
by: Liu, Jinmei, et al.
Published: (2024)
PAGE: Domain-Incremental Adaptation with Past-Agnostic Generative Replay for Smart Healthcare
by: Li, Chia-Hao, et al.
Published: (2024)
by: Li, Chia-Hao, et al.
Published: (2024)
Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay
by: Binici, Kuluhan, et al.
Published: (2022)
by: Binici, Kuluhan, et al.
Published: (2022)
Influential Bandits: Pulling an Arm May Change the Environment
by: Sato, Ryoma, et al.
Published: (2025)
by: Sato, Ryoma, et al.
Published: (2025)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
by: Mahrooghi, Ilia, et al.
Published: (2026)
by: Mahrooghi, Ilia, et al.
Published: (2026)
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach
by: Li, Wenyun, et al.
Published: (2025)
by: Li, Wenyun, et al.
Published: (2025)
D-SPEAR: Dual-Stream Prioritized Experience Adaptive Replay for Stable Reinforcement Learning in Robotic Manipulation
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Similar Items
-
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
by: Romio, Gabriel, et al.
Published: (2026) -
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
by: Cho, Minjae, et al.
Published: (2026) -
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026) -
Experience Replay with Random Reshuffling
by: Fujita, Yasuhiro
Published: (2025) -
Sample Efficient Experience Replay in Non-stationary Environments
by: Duan, Tianyang, et al.
Published: (2025)