Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Goodall, Alexander W., Court, Edwin Hamel-De le, Belardinelli, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probabilistic Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
ProSh: Probabilistic Shielding for Model-free Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
by: Court, Edwin Hamel-De le, et al.
Published: (2025)
Robust Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2026)
by: Court, Edwin Hamel-De le, et al.
Published: (2026)
Safe Reinforcement Learning via Recovery-based Shielding with Gaussian Process Dynamics Models
by: Goodall, Alexander W., et al.
Published: (2026)
by: Goodall, Alexander W., et al.
Published: (2026)
Approximate Model-Based Shielding for Safe Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2023)
by: Goodall, Alexander W., et al.
Published: (2023)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
Leveraging Approximate Model-based Shielding for Probabilistic Safety Guarantees in Continuous Environments
by: Goodall, Alexander W., et al.
Published: (2024)
by: Goodall, Alexander W., et al.
Published: (2024)
Synthesis of Safety Specifications for Probabilistic Systems
by: Ohlmann, Gaspard, et al.
Published: (2025)
by: Ohlmann, Gaspard, et al.
Published: (2025)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement
by: Liang, Haodong, et al.
Published: (2026)
by: Liang, Haodong, et al.
Published: (2026)
Search-Based Adversarial Estimates for Improving Sample Efficiency in Off-Policy Reinforcement Learning
by: Malato, Federico, et al.
Published: (2025)
by: Malato, Federico, et al.
Published: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning
by: Zhuang, Yuan, et al.
Published: (2026)
by: Zhuang, Yuan, et al.
Published: (2026)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Automated Off-Policy Estimator Selection via Supervised Learning
by: Felicioni, Nicolò, et al.
Published: (2024)
by: Felicioni, Nicolò, et al.
Published: (2024)
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026)
by: Yu, Tianrun, et al.
Published: (2026)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Beyond Expected Return: Accounting for Policy Reproducibility when Evaluating Reinforcement Learning Algorithms
by: Flageat, Manon, et al.
Published: (2023)
by: Flageat, Manon, et al.
Published: (2023)
Zero-Shot Off-Policy Learning
by: Asadulaev, Arip, et al.
Published: (2026)
by: Asadulaev, Arip, et al.
Published: (2026)
A Variance-Reduced Cubic-Regularized Newton for Policy Optimization
by: Sun, Cheng, et al.
Published: (2025)
by: Sun, Cheng, et al.
Published: (2025)
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
by: Wang, Ruhan, et al.
Published: (2024)
by: Wang, Ruhan, et al.
Published: (2024)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Moments Matter:Stabilizing Policy Optimization using Return Distributions
by: Jabs, Dennis, et al.
Published: (2026)
by: Jabs, Dennis, et al.
Published: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning
by: Farhi, Nadir
Published: (2025)
by: Farhi, Nadir
Published: (2025)
On the Structural Non-Preservation of Epistemic Behaviour under Policy Transformation
by: Galozy, Alexander
Published: (2026)
by: Galozy, Alexander
Published: (2026)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Multi-objective Reinforcement Learning with Nonlinear Preferences: Provable Approximation for Maximizing Expected Scalarized Return
by: Peng, Nianli, et al.
Published: (2023)
by: Peng, Nianli, et al.
Published: (2023)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
Model-Based Epistemic Variance of Values for Risk-Aware Policy Optimization
by: Luis, Carlos E., et al.
Published: (2023)
by: Luis, Carlos E., et al.
Published: (2023)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
Similar Items
-
Probabilistic Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025) -
ProSh: Probabilistic Shielding for Model-free Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2025) -
Robust Shielding for Safe Reinforcement Learning
by: Court, Edwin Hamel-De le, et al.
Published: (2026) -
Safe Reinforcement Learning via Recovery-based Shielding with Gaussian Process Dynamics Models
by: Goodall, Alexander W., et al.
Published: (2026) -
Approximate Model-Based Shielding for Safe Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2023)