Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
Fuente:
arXiv
Saved in:
| Main Authors: | Mead, Harry, Costen, Clarissa, Lacerda, Bruno, Hawes, Nick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Insurance Reserving with CVaR-Constrained Reinforcement Learning under Macroeconomic Regimes
by: Dong, Stella C.
Published: (2025)
by: Dong, Stella C.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Boosting CVaR Policy Optimization with Quantile Gradients
by: Luo, Yudong, et al.
Published: (2026)
by: Luo, Yudong, et al.
Published: (2026)
Improving Regret Approximation for Unsupervised Dynamic Environment Generation
by: Mead, Harry, et al.
Published: (2026)
by: Mead, Harry, et al.
Published: (2026)
Monte Carlo Tree Search with Boltzmann Exploration
by: Painter, Michael, et al.
Published: (2024)
by: Painter, Michael, et al.
Published: (2024)
JaxWildfire: A GPU-Accelerated Wildfire Simulator for Reinforcement Learning
by: Çakır, Ufuk, et al.
Published: (2025)
by: Çakır, Ufuk, et al.
Published: (2025)
Distributionally Robust Safety Verification of Neural Networks via Worst-Case CVaR
by: Kishida, Masako
Published: (2025)
by: Kishida, Masako
Published: (2025)
Tackling GNARLy Problems: Graph Neural Algorithmic Reasoning Reimagined through Reinforcement Learning
by: Schutz, Alex, et al.
Published: (2025)
by: Schutz, Alex, et al.
Published: (2025)
CVaR-Based Variational Quantum Optimization for User Association in Handoff-Aware Vehicular Networks
by: Yan, Zijiang, et al.
Published: (2025)
by: Yan, Zijiang, et al.
Published: (2025)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
by: Rutherford, Alexander, et al.
Published: (2024)
by: Rutherford, Alexander, et al.
Published: (2024)
A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization
by: Luo, Yudong, et al.
Published: (2024)
by: Luo, Yudong, et al.
Published: (2024)
Actor-Critic Algorithm for Dynamic Expectile and CVaR
by: Luo, Yudong, et al.
Published: (2026)
by: Luo, Yudong, et al.
Published: (2026)
ClauseLens: Clause-Grounded, CVaR-Constrained Reinforcement Learning for Trustworthy Reinsurance Pricing
by: Dong, Stella C., et al.
Published: (2025)
by: Dong, Stella C., et al.
Published: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Near-Optimal Sample Complexity for Iterated CVaR Reinforcement Learning with a Generative Model
by: Deng, Zilong, et al.
Published: (2025)
by: Deng, Zilong, et al.
Published: (2025)
Autonomous Sparse Mean-CVaR Portfolio Optimization
by: Lin, Yizun, et al.
Published: (2024)
by: Lin, Yizun, et al.
Published: (2024)
A Finite-State Controller Based Offline Solver for Deterministic POMDPs
by: Schutz, Alex, et al.
Published: (2025)
by: Schutz, Alex, et al.
Published: (2025)
Holder Policy Optimisation
by: Chen, Yuxiang, et al.
Published: (2026)
by: Chen, Yuxiang, et al.
Published: (2026)
Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds
by: Pariente, Yaacov, et al.
Published: (2026)
by: Pariente, Yaacov, et al.
Published: (2026)
Online Risk-Averse Planning in POMDPs Using Iterated CVaR Value Function
by: Pariente, Yaacov, et al.
Published: (2026)
by: Pariente, Yaacov, et al.
Published: (2026)
On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes
by: Godbout, Mathieu, et al.
Published: (2025)
by: Godbout, Mathieu, et al.
Published: (2025)
The Privacy Price of Tail-Risk Learning: Effective Tail Sample Size in Differentially Private CVaR Optimization
by: Mansouri, El Mustapha
Published: (2026)
by: Mansouri, El Mustapha
Published: (2026)
DITTO: Offline Imitation Learning with World Models
by: DeMoss, Branton, et al.
Published: (2023)
by: DeMoss, Branton, et al.
Published: (2023)
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024)
by: Macfarlane, Matthew V, et al.
Published: (2024)
Instantiating Bayesian CVaR lower bounds in Interactive Decision Making Problems
by: Bongole, Raghav, et al.
Published: (2026)
by: Bongole, Raghav, et al.
Published: (2026)
Mirror Learning: A Unifying Framework of Policy Optimisation
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
by: Kuba, Jakub Grudzien, et al.
Published: (2022)
Learning Optimal and Sample-Efficient Decision Policies with Guarantees
by: Shao, Daqian
Published: (2026)
by: Shao, Daqian
Published: (2026)
Multi-Agent Regime-Conditioned Diffusion (MARCD) for CVaR-Constrained Portfolio Decisions
by: Alzahrani, Ali Atiah
Published: (2025)
by: Alzahrani, Ali Atiah
Published: (2025)
Statistical Robustness of Interval CVaR Based Regression Models under Perturbation and Contamination
by: You, Yulei, et al.
Published: (2026)
by: You, Yulei, et al.
Published: (2026)
Moments Matter:Stabilizing Policy Optimization using Return Distributions
by: Jabs, Dennis, et al.
Published: (2026)
by: Jabs, Dennis, et al.
Published: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
by: Miao, Yuchun, et al.
Published: (2026)
by: Miao, Yuchun, et al.
Published: (2026)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
Gradients: When Markets Meet Fine-tuning -- A Distributed Approach to Model Optimisation
by: Subia-Waud, Christopher
Published: (2025)
by: Subia-Waud, Christopher
Published: (2025)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
by: Sienkiewicz, Bruno, et al.
Published: (2026)
by: Sienkiewicz, Bruno, et al.
Published: (2026)
R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning
by: Farhi, Nadir
Published: (2025)
by: Farhi, Nadir
Published: (2025)
Gaussian Process Aggregation for Root-Parallel Monte Carlo Tree Search with Continuous Actions
by: Xiao, Junlin, et al.
Published: (2025)
by: Xiao, Junlin, et al.
Published: (2025)
Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk
by: Woydt, Tim, et al.
Published: (2026)
by: Woydt, Tim, et al.
Published: (2026)
Similar Items
-
Adaptive Insurance Reserving with CVaR-Constrained Reinforcement Learning under Macroeconomic Regimes
by: Dong, Stella C.
Published: (2025) -
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026) -
Boosting CVaR Policy Optimization with Quantile Gradients
by: Luo, Yudong, et al.
Published: (2026) -
Improving Regret Approximation for Unsupervised Dynamic Environment Generation
by: Mead, Harry, et al.
Published: (2026) -
Monte Carlo Tree Search with Boltzmann Exploration
by: Painter, Michael, et al.
Published: (2024)