When Do Off-Policy and On-Policy Policy Gradient Methods Align?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mambelli, Davide, Bongers, Stephan, Zoeter, Onno, Spaan, Matthijs T. J., Oliehoek, Frans A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation
von: Brita, Catalin E., et al.
Veröffentlicht: (2024)
von: Brita, Catalin E., et al.
Veröffentlicht: (2024)
Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization
von: Lee, Wook, et al.
Veröffentlicht: (2025)
von: Lee, Wook, et al.
Veröffentlicht: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
Difference Rewards Policy Gradients
von: Castellini, Jacopo, et al.
Veröffentlicht: (2020)
von: Castellini, Jacopo, et al.
Veröffentlicht: (2020)
Trust-Region Twisted Policy Improvement
von: de Vries, Joery A., et al.
Veröffentlicht: (2025)
von: de Vries, Joery A., et al.
Veröffentlicht: (2025)
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
von: Suau, Miguel, et al.
Veröffentlicht: (2022)
von: Suau, Miguel, et al.
Veröffentlicht: (2022)
Sparse Masked Attention Policies for Reliable Generalization
von: Horsch, Caroline, et al.
Veröffentlicht: (2026)
von: Horsch, Caroline, et al.
Veröffentlicht: (2026)
Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response
von: Bighashdel, Ariyan, et al.
Veröffentlicht: (2026)
von: Bighashdel, Ariyan, et al.
Veröffentlicht: (2026)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
von: Weltevrede, Max, et al.
Veröffentlicht: (2025)
von: Weltevrede, Max, et al.
Veröffentlicht: (2025)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
von: Meng, Wenjia, et al.
Veröffentlicht: (2024)
Navigating Trade-offs: Policy Summarization for Multi-Objective Reinforcement Learning
von: Osika, Zuzanna, et al.
Veröffentlicht: (2024)
von: Osika, Zuzanna, et al.
Veröffentlicht: (2024)
Evaluating and Correcting Performative Effects of Decision Support Systems via Causal Domain Shift
von: Boeken, Philip, et al.
Veröffentlicht: (2024)
von: Boeken, Philip, et al.
Veröffentlicht: (2024)
Unidentified and Confounded? Understanding Two-Tower Models for Unbiased Learning to Rank
von: Hager, Philipp, et al.
Veröffentlicht: (2025)
von: Hager, Philipp, et al.
Veröffentlicht: (2025)
Conditional Forecasts and Proper Scoring Rules for Reliable and Accurate Performative Predictions
von: Boeken, Philip, et al.
Veröffentlicht: (2025)
von: Boeken, Philip, et al.
Veröffentlicht: (2025)
Learning General Policies with Policy Gradient Methods
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2025)
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2025)
Exploring Equity of Climate Policies using Multi-Agent Multi-Objective Reinforcement Learning
von: Biswas, Palok, et al.
Veröffentlicht: (2025)
von: Biswas, Palok, et al.
Veröffentlicht: (2025)
CLAX: Fast and Flexible Neural Click Models in JAX
von: Hager, Philipp, et al.
Veröffentlicht: (2025)
von: Hager, Philipp, et al.
Veröffentlicht: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
von: Tanaka, Koichi, et al.
Veröffentlicht: (2026)
von: Tanaka, Koichi, et al.
Veröffentlicht: (2026)
Foundations of Structural Causal Models with Latent Selection
von: Chen, Leihao, et al.
Veröffentlicht: (2024)
von: Chen, Leihao, et al.
Veröffentlicht: (2024)
Policy Gradient Methods for Distortion Risk Measures
von: Vijayan, Nithia, et al.
Veröffentlicht: (2021)
von: Vijayan, Nithia, et al.
Veröffentlicht: (2021)
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
von: Mroueh, Youssef, et al.
Veröffentlicht: (2025)
von: Mroueh, Youssef, et al.
Veröffentlicht: (2025)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
von: Onoda, Ku, et al.
Veröffentlicht: (2026)
von: Onoda, Ku, et al.
Veröffentlicht: (2026)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
Unifying On- and Off-Policy Variance Reduction Methods
von: Jeunen, Olivier
Veröffentlicht: (2026)
von: Jeunen, Olivier
Veröffentlicht: (2026)
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
von: Xia, Linxuan, et al.
Veröffentlicht: (2026)
von: Xia, Linxuan, et al.
Veröffentlicht: (2026)
Elementary Analysis of Policy Gradient Methods
von: Liu, Jiacai, et al.
Veröffentlicht: (2024)
von: Liu, Jiacai, et al.
Veröffentlicht: (2024)
Reevaluating Policy Gradient Methods for Imperfect-Information Games
von: Rudolph, Max, et al.
Veröffentlicht: (2025)
von: Rudolph, Max, et al.
Veröffentlicht: (2025)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
von: Ankile, Lars, et al.
Veröffentlicht: (2025)
von: Ankile, Lars, et al.
Veröffentlicht: (2025)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2024)
von: Kiyohara, Haruka, et al.
Veröffentlicht: (2024)
Deep Gaussian Process Proximal Policy Optimization
von: van der Lende, Matthijs, et al.
Veröffentlicht: (2025)
von: van der Lende, Matthijs, et al.
Veröffentlicht: (2025)
Mollification Effects of Policy Gradient Methods
von: Wang, Tao, et al.
Veröffentlicht: (2024)
von: Wang, Tao, et al.
Veröffentlicht: (2024)
Learning Optimal Deterministic Policies with Stochastic Policy Gradients
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients
von: DeWeese, Alex, et al.
Veröffentlicht: (2026)
von: DeWeese, Alex, et al.
Veröffentlicht: (2026)
Group Policy Gradient
von: Chen, Junhua, et al.
Veröffentlicht: (2025)
von: Chen, Junhua, et al.
Veröffentlicht: (2025)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
von: Murgoci, Vlad, et al.
Veröffentlicht: (2026)
von: Murgoci, Vlad, et al.
Veröffentlicht: (2026)
Scaling Internal-State Policy-Gradient Methods for POMDPs
von: Aberdeen, Douglas, et al.
Veröffentlicht: (2025)
von: Aberdeen, Douglas, et al.
Veröffentlicht: (2025)
Cross-Validated Off-Policy Evaluation
von: Cief, Matej, et al.
Veröffentlicht: (2024)
von: Cief, Matej, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
von: Suau, Miguel, et al.
Veröffentlicht: (2023) -
SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation
von: Brita, Catalin E., et al.
Veröffentlicht: (2024) -
Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization
von: Lee, Wook, et al.
Veröffentlicht: (2025) -
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
von: Li, Guopeng, et al.
Veröffentlicht: (2026) -
Difference Rewards Policy Gradients
von: Castellini, Jacopo, et al.
Veröffentlicht: (2020)