Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenglin, Ruan, Guangchun, Geng, Hua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909431737876480
author Li, Chenglin
Ruan, Guangchun
Geng, Hua
author_facet Li, Chenglin
Ruan, Guangchun
Geng, Hua
contents Safe reinforcement learning (RL) is a popular and versatile paradigm to learn reward-maximizing policies with safety guarantees. Previous works tend to express the safety constraints in an expectation form due to the ease of implementation, but this turns out to be ineffective in maintaining safety constraints with high probability. To this end, we move to the quantile-constrained RL that enables a higher level of safety without any expectation-form approximations. We directly estimate the quantile gradients through sampling and provide the theoretical proofs of convergence. Then a tilted update strategy for quantile gradients is implemented to compensate the asymmetric distributional density, with a direct benefit of return performance. Experiments demonstrate that the proposed model fully meets safety requirements (quantile constraints) while outperforming the state-of-the-art benchmarks with higher return.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning
Li, Chenglin
Ruan, Guangchun
Geng, Hua
Machine Learning
Artificial Intelligence
Safe reinforcement learning (RL) is a popular and versatile paradigm to learn reward-maximizing policies with safety guarantees. Previous works tend to express the safety constraints in an expectation form due to the ease of implementation, but this turns out to be ineffective in maintaining safety constraints with high probability. To this end, we move to the quantile-constrained RL that enables a higher level of safety without any expectation-form approximations. We directly estimate the quantile gradients through sampling and provide the theoretical proofs of convergence. Then a tilted update strategy for quantile gradients is implemented to compensate the asymmetric distributional density, with a direct benefit of return performance. Experiments demonstrate that the proposed model fully meets safety requirements (quantile constraints) while outperforming the state-of-the-art benchmarks with higher return.
title Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2412.13184