On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Hua, Xie, Wei, Feng, M. Ben |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variance Reduction Based Experience Replay for Policy Optimization
by: Zheng, Hua, et al.
Published: (2026)
by: Zheng, Hua, et al.
Published: (2026)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
by: Lim, Han-Dong, et al.
Published: (2023)
by: Lim, Han-Dong, et al.
Published: (2023)
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
by: Ospanov, Azim, et al.
Published: (2024)
by: Ospanov, Azim, et al.
Published: (2024)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
Experience Replay with Random Reshuffling
by: Fujita, Yasuhiro
Published: (2025)
by: Fujita, Yasuhiro
Published: (2025)
ROER: Regularized Optimal Experience Replay
by: Li, Changling, et al.
Published: (2024)
by: Li, Changling, et al.
Published: (2024)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
RePO: Replay-Enhanced Policy Optimization
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
by: Yenicesu, Arda Sarp, et al.
Published: (2024)
by: Yenicesu, Arda Sarp, et al.
Published: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
by: Sun, Cheng, et al.
Published: (2025)
by: Sun, Cheng, et al.
Published: (2025)
Mastering the Game of Go with Self-play Experience Replay
by: Liu, Jingbin, et al.
Published: (2026)
by: Liu, Jingbin, et al.
Published: (2026)
LTL-Constrained Policy Optimization with Cycle Experience Replay
by: Shah, Ameesh, et al.
Published: (2024)
by: Shah, Ameesh, et al.
Published: (2024)
Second-Order Convergence in Private Stochastic Non-Convex Optimization
by: Tao, Youming, et al.
Published: (2025)
by: Tao, Youming, et al.
Published: (2025)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
by: Luis, Carlos E., et al.
Published: (2023)
by: Luis, Carlos E., et al.
Published: (2023)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
by: Zhao, Kaiyan, et al.
Published: (2024)
by: Zhao, Kaiyan, et al.
Published: (2024)
Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
by: Lyu, Guwen, et al.
Published: (2024)
by: Lyu, Guwen, et al.
Published: (2024)
Epistemic Error Decomposition for Multi-step Time Series Forecasting: Rethinking Bias-Variance in Recursive and Direct Strategies
by: Green, Riku, et al.
Published: (2025)
by: Green, Riku, et al.
Published: (2025)
On the Rate of Convergence of Kolmogorov-Arnold Network Regression Estimators
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
by: Balasubramanian, Krishnakumar
Published: (2026)
by: Balasubramanian, Krishnakumar
Published: (2026)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
by: Pan, Chengjun, et al.
Published: (2026)
by: Pan, Chengjun, et al.
Published: (2026)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
by: Liu, Jinyi, et al.
Published: (2023)
by: Liu, Jinyi, et al.
Published: (2023)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP
by: Sangadi, Tejaram, et al.
Published: (2024)
by: Sangadi, Tejaram, et al.
Published: (2024)
Defense without Forgetting: Continual Adversarial Defense with Anisotropic & Isotropic Pseudo Replay
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
Sample Efficient Experience Replay in Non-stationary Environments
by: Duan, Tianyang, et al.
Published: (2025)
by: Duan, Tianyang, et al.
Published: (2025)
Representation Convergence: Mutual Distillation is Secretly a Form of Regularization
by: Xie, Zhengpeng, et al.
Published: (2025)
by: Xie, Zhengpeng, et al.
Published: (2025)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
by: Romio, Gabriel, et al.
Published: (2026)
by: Romio, Gabriel, et al.
Published: (2026)
Theoretical Convergence of SMOTE-Generated Samples
by: Kamalov, Firuz, et al.
Published: (2026)
by: Kamalov, Firuz, et al.
Published: (2026)
Convergent Linear Representations of Emergent Misalignment
by: Soligo, Anna, et al.
Published: (2025)
by: Soligo, Anna, et al.
Published: (2025)
Convergent World Representations and Divergent Tasks
by: Park, Core Francisco
Published: (2026)
by: Park, Core Francisco
Published: (2026)
On the Convergence of Continual Learning with Adaptive Methods
by: Han, Seungyub, et al.
Published: (2024)
by: Han, Seungyub, et al.
Published: (2024)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
CIER: A Novel Experience Replay Approach with Causal Inference in Deep Reinforcement Learning
by: Wang, Jingwen, et al.
Published: (2024)
by: Wang, Jingwen, et al.
Published: (2024)
Similar Items
-
Variance Reduction Based Experience Replay for Policy Optimization
by: Zheng, Hua, et al.
Published: (2026) -
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
by: Lim, Han-Dong, et al.
Published: (2023) -
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
by: Ospanov, Azim, et al.
Published: (2024) -
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026) -
Experience Replay with Random Reshuffling
by: Fujita, Yasuhiro
Published: (2025)