ExO-PPO: an Extended Off-policy Proximal Policy Optimization Algorithm
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hanyong, Yang, Menglong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CIM-PPO:Proximal Policy Optimization with Liu-Correntropy Induced Metric
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021)
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021)
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization
von: Sane, Soham
Veröffentlicht: (2025)
von: Sane, Soham
Veröffentlicht: (2025)
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)
ESPO: Early-Stopping Proximal Policy Optimization
von: Li, Zihang, et al.
Veröffentlicht: (2026)
von: Li, Zihang, et al.
Veröffentlicht: (2026)
Complexity-Regularized Proximal Policy Optimization
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
von: Serfilippi, Luca, et al.
Veröffentlicht: (2025)
Beyond the Boundaries of Proximal Policy Optimization
von: Tan, Charlie B., et al.
Veröffentlicht: (2024)
von: Tan, Charlie B., et al.
Veröffentlicht: (2024)
Proximal Policy Optimization with Adaptive Exploration
von: Lixandru, Andrei
Veröffentlicht: (2024)
von: Lixandru, Andrei
Veröffentlicht: (2024)
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
KIPPO: Koopman-Inspired Proximal Policy Optimization
von: Cozma, Andrei, et al.
Veröffentlicht: (2025)
von: Cozma, Andrei, et al.
Veröffentlicht: (2025)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
von: Liu, Pai, et al.
Veröffentlicht: (2025)
von: Liu, Pai, et al.
Veröffentlicht: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
von: Liu, Jiashun, et al.
Veröffentlicht: (2025)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
Learning Branching Policies for MILPs with Proximal Policy Optimization
von: Mhamed, Abdelouahed Ben, et al.
Veröffentlicht: (2025)
von: Mhamed, Abdelouahed Ben, et al.
Veröffentlicht: (2025)
Proximal Policy Distillation
von: Spigler, Giacomo
Veröffentlicht: (2024)
von: Spigler, Giacomo
Veröffentlicht: (2024)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
Transductive Off-policy Proximal Policy Optimization
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
von: Gan, Yaozhong, et al.
Veröffentlicht: (2024)
A dynamical clipping approach with task feedback for Proximal Policy Optimization
von: Zhang, Ziqi, et al.
Veröffentlicht: (2023)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2023)
Efficient Deep Reinforcement Learning with Predictive Processing Proximal Policy Optimization
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2022)
von: Küçükoğlu, Burcu, et al.
Veröffentlicht: (2022)
ExOSITO: Explainable Off-Policy Learning with Side Information for Intensive Care Unit Blood Test Orders
von: Ji, Zongliang, et al.
Veröffentlicht: (2025)
von: Ji, Zongliang, et al.
Veröffentlicht: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2025)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
Bootstrap Off-policy with World Model
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
CUER: Corrected Uniform Experience Replay for Off-Policy Continuous Deep Reinforcement Learning Algorithms
von: Yenicesu, Arda Sarp, et al.
Veröffentlicht: (2024)
von: Yenicesu, Arda Sarp, et al.
Veröffentlicht: (2024)
Primal-Dual Spectral Representation for Off-policy Evaluation
von: Hu, Yang, et al.
Veröffentlicht: (2024)
von: Hu, Yang, et al.
Veröffentlicht: (2024)
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
Zero-Shot Off-Policy Learning
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
von: Asadulaev, Arip, et al.
Veröffentlicht: (2026)
Clustering Context in Off-Policy Evaluation
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
von: Guzman-Olivares, Daniel, et al.
Veröffentlicht: (2025)
Concept-driven Off Policy Evaluation
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
von: Majumdar, Ritam, et al.
Veröffentlicht: (2024)
Mode-Dependent Rectification for Stable PPO Training
von: Mohamad, Mohamad, et al.
Veröffentlicht: (2026)
von: Mohamad, Mohamad, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CIM-PPO:Proximal Policy Optimization with Liu-Correntropy Induced Metric
von: Guo, Yunxiao, et al.
Veröffentlicht: (2021) -
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization
von: Sane, Soham
Veröffentlicht: (2025) -
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025) -
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2026)