Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mead, Harry, Costen, Clarissa, Lacerda, Bruno, Hawes, Nick
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908456547516416
author Mead, Harry
Costen, Clarissa
Lacerda, Bruno
Hawes, Nick
author_facet Mead, Harry
Costen, Clarissa
Lacerda, Bruno
Hawes, Nick
contents When optimising for conditional value at risk (CVaR) using policy gradients (PG), current methods rely on discarding a large proportion of trajectories, resulting in poor sample efficiency. We propose a reformulation of the CVaR optimisation problem by capping the total return of trajectories used in training, rather than simply discarding them, and show that this is equivalent to the original problem if the cap is set appropriately. We show, with empirical results in an number of environments, that this reformulation of the problem results in consistently improved performance compared to baselines. We have made all our code available here: https://github.com/HarryMJMead/cvar-return-capping.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
Mead, Harry
Costen, Clarissa
Lacerda, Bruno
Hawes, Nick
Machine Learning
Artificial Intelligence
When optimising for conditional value at risk (CVaR) using policy gradients (PG), current methods rely on discarding a large proportion of trajectories, resulting in poor sample efficiency. We propose a reformulation of the CVaR optimisation problem by capping the total return of trajectories used in training, rather than simply discarding them, and show that this is equivalent to the original problem if the cap is set appropriately. We show, with empirical results in an number of environments, that this reformulation of the problem results in consistently improved performance compared to baselines. We have made all our code available here: https://github.com/HarryMJMead/cvar-return-capping.
title Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.20887