Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: McCarthy, James, Marinescu, Radu, Daly, Elizabeth, Dusparic, Ivana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911124085014528
author McCarthy, James
Marinescu, Radu
Daly, Elizabeth
Dusparic, Ivana
author_facet McCarthy, James
Marinescu, Radu
Daly, Elizabeth
Dusparic, Ivana
contents Risk-averse Constrained Reinforcement Learning (RaCRL) aims to learn policies that minimise the likelihood of rare and catastrophic constraint violations caused by an environment's inherent randomness. In general, risk-aversion leads to conservative exploration of the environment which typically results in converging to sub-optimal policies that fail to adequately maximise reward or, in some cases, fail to achieve the goal. In this paper, we propose an exploration-based approach for RaCRL called Optimistic Risk-averse Actor Critic (ORAC), which constructs an exploratory policy by maximising a local upper confidence bound of the state-action reward value function whilst minimising a local lower confidence bound of the risk-averse state-action cost value function. Specifically, at each step, the weighting assigned to the cost value is increased or decreased if it exceeds or falls below the safety constraint value. This way the policy is encouraged to explore uncertain regions of the environment to discover high reward states whilst still satisfying the safety constraints. Our experimental results demonstrate that the ORAC approach prevents convergence to sub-optimal policies and improves significantly the reward-cost trade-off in various continuous control tasks such as Safety-Gymnasium and a complex building energy management environment CityLearn.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
McCarthy, James
Marinescu, Radu
Daly, Elizabeth
Dusparic, Ivana
Machine Learning
Artificial Intelligence
Risk-averse Constrained Reinforcement Learning (RaCRL) aims to learn policies that minimise the likelihood of rare and catastrophic constraint violations caused by an environment's inherent randomness. In general, risk-aversion leads to conservative exploration of the environment which typically results in converging to sub-optimal policies that fail to adequately maximise reward or, in some cases, fail to achieve the goal. In this paper, we propose an exploration-based approach for RaCRL called Optimistic Risk-averse Actor Critic (ORAC), which constructs an exploratory policy by maximising a local upper confidence bound of the state-action reward value function whilst minimising a local lower confidence bound of the risk-averse state-action cost value function. Specifically, at each step, the weighting assigned to the cost value is increased or decreased if it exceeds or falls below the safety constraint value. This way the policy is encouraged to explore uncertain regions of the environment to discover high reward states whilst still satisfying the safety constraints. Our experimental results demonstrate that the ORAC approach prevents convergence to sub-optimal policies and improves significantly the reward-cost trade-off in various continuous control tasks such as Safety-Gymnasium and a complex building energy management environment CityLearn.
title Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.08793