Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stolz, Roland, Eichelbeck, Michael, Althoff, Matthias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914173358702592
author Stolz, Roland
Eichelbeck, Michael
Althoff, Matthias
author_facet Stolz, Roland
Eichelbeck, Michael
Althoff, Matthias
contents In reinforcement learning (RL), it is often advantageous to consider additional constraints on the action space to ensure safety or action relevance. Existing work on such action-constrained RL faces challenges regarding effective policy updates, computational efficiency, and predictable runtime. Recent work proposes to use truncated normal distributions for stochastic policy gradient methods. However, the computation of key characteristics, such as the entropy, log-probability, and their gradients, becomes intractable under complex constraints. Hence, prior work approximates these using the non-truncated distributions, which severely degrades performance. We argue that accurate estimation of these characteristics is crucial in the action-constrained RL setting, and propose efficient numerical approximations for them. We also provide an efficient sampling strategy for truncated policy distributions and validate our approach on three benchmark environments, which demonstrate significant performance improvements when using accurate estimations.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions
Stolz, Roland
Eichelbeck, Michael
Althoff, Matthias
Machine Learning
Systems and Control
In reinforcement learning (RL), it is often advantageous to consider additional constraints on the action space to ensure safety or action relevance. Existing work on such action-constrained RL faces challenges regarding effective policy updates, computational efficiency, and predictable runtime. Recent work proposes to use truncated normal distributions for stochastic policy gradient methods. However, the computation of key characteristics, such as the entropy, log-probability, and their gradients, becomes intractable under complex constraints. Hence, prior work approximates these using the non-truncated distributions, which severely degrades performance. We argue that accurate estimation of these characteristics is crucial in the action-constrained RL setting, and propose efficient numerical approximations for them. We also provide an efficient sampling strategy for truncated policy distributions and validate our approach on three benchmark environments, which demonstrate significant performance improvements when using accurate estimations.
title Improving Stochastic Action-Constrained Reinforcement Learning via Truncated Distributions
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2511.22406