On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Hua, Xie, Wei, Feng, M. Ben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Variance Reduction Based Experience Replay for Policy Optimization
von: Zheng, Hua, et al.
Veröffentlicht: (2026)
von: Zheng, Hua, et al.
Veröffentlicht: (2026)
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023)
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
VLM-Guided Experience Replay
von: Sharony, Elad, et al.
Veröffentlicht: (2026)
von: Sharony, Elad, et al.
Veröffentlicht: (2026)
Experience Replay with Random Reshuffling
von: Fujita, Yasuhiro
Veröffentlicht: (2025)
von: Fujita, Yasuhiro
Veröffentlicht: (2025)
ROER: Regularized Optimal Experience Replay
von: Li, Changling, et al.
Veröffentlicht: (2024)
von: Li, Changling, et al.
Veröffentlicht: (2024)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
von: Perkins, Daniel, et al.
Veröffentlicht: (2025)
von: Perkins, Daniel, et al.
Veröffentlicht: (2025)
RePO: Replay-Enhanced Policy Optimization
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
von: Yenicesu, Arda Sarp, et al.
Veröffentlicht: (2024)
von: Yenicesu, Arda Sarp, et al.
Veröffentlicht: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
von: Asad, Reza, et al.
Veröffentlicht: (2024)
von: Asad, Reza, et al.
Veröffentlicht: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
von: Mondal, Washim Uddin, et al.
Veröffentlicht: (2024)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
von: Sun, Cheng, et al.
Veröffentlicht: (2025)
von: Sun, Cheng, et al.
Veröffentlicht: (2025)
Mastering the Game of Go with Self-play Experience Replay
von: Liu, Jingbin, et al.
Veröffentlicht: (2026)
von: Liu, Jingbin, et al.
Veröffentlicht: (2026)
LTL-Constrained Policy Optimization with Cycle Experience Replay
von: Shah, Ameesh, et al.
Veröffentlicht: (2024)
von: Shah, Ameesh, et al.
Veröffentlicht: (2024)
Second-Order Convergence in Private Stochastic Non-Convex Optimization
von: Tao, Youming, et al.
Veröffentlicht: (2025)
von: Tao, Youming, et al.
Veröffentlicht: (2025)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
von: Luo, Yu, et al.
Veröffentlicht: (2026)
von: Luo, Yu, et al.
Veröffentlicht: (2026)
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
von: Luis, Carlos E., et al.
Veröffentlicht: (2023)
von: Luis, Carlos E., et al.
Veröffentlicht: (2023)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024)
Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
von: Lyu, Guwen, et al.
Veröffentlicht: (2024)
von: Lyu, Guwen, et al.
Veröffentlicht: (2024)
Epistemic Error Decomposition for Multi-step Time Series Forecasting: Rethinking Bias-Variance in Recursive and Direct Strategies
von: Green, Riku, et al.
Veröffentlicht: (2025)
von: Green, Riku, et al.
Veröffentlicht: (2025)
On the Rate of Convergence of Kolmogorov-Arnold Network Regression Estimators
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
von: Balasubramanian, Krishnakumar
Veröffentlicht: (2026)
von: Balasubramanian, Krishnakumar
Veröffentlicht: (2026)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
von: Liu, Jinyi, et al.
Veröffentlicht: (2023)
von: Liu, Jinyi, et al.
Veröffentlicht: (2023)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP
von: Sangadi, Tejaram, et al.
Veröffentlicht: (2024)
von: Sangadi, Tejaram, et al.
Veröffentlicht: (2024)
Defense without Forgetting: Continual Adversarial Defense with Anisotropic & Isotropic Pseudo Replay
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
Sample Efficient Experience Replay in Non-stationary Environments
von: Duan, Tianyang, et al.
Veröffentlicht: (2025)
von: Duan, Tianyang, et al.
Veröffentlicht: (2025)
Representation Convergence: Mutual Distillation is Secretly a Form of Regularization
von: Xie, Zhengpeng, et al.
Veröffentlicht: (2025)
von: Xie, Zhengpeng, et al.
Veröffentlicht: (2025)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
Theoretical Convergence of SMOTE-Generated Samples
von: Kamalov, Firuz, et al.
Veröffentlicht: (2026)
von: Kamalov, Firuz, et al.
Veröffentlicht: (2026)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Convergent World Representations and Divergent Tasks
von: Park, Core Francisco
Veröffentlicht: (2026)
von: Park, Core Francisco
Veröffentlicht: (2026)
On the Convergence of Continual Learning with Adaptive Methods
von: Han, Seungyub, et al.
Veröffentlicht: (2024)
von: Han, Seungyub, et al.
Veröffentlicht: (2024)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
von: Liu, Xinyu, et al.
Veröffentlicht: (2025)
CIER: A Novel Experience Replay Approach with Causal Inference in Deep Reinforcement Learning
von: Wang, Jingwen, et al.
Veröffentlicht: (2024)
von: Wang, Jingwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Variance Reduction Based Experience Replay for Policy Optimization
von: Zheng, Hua, et al.
Veröffentlicht: (2026) -
Finite-Time Analysis of Temporal Difference Learning with Experience Replay
von: Lim, Han-Dong, et al.
Veröffentlicht: (2023) -
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
von: Ospanov, Azim, et al.
Veröffentlicht: (2024) -
VLM-Guided Experience Replay
von: Sharony, Elad, et al.
Veröffentlicht: (2026) -
Experience Replay with Random Reshuffling
von: Fujita, Yasuhiro
Veröffentlicht: (2025)