Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yuxiao, Sinha, Arunesh, Varakantham, Pradeep
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910290392645632
author Lu, Yuxiao
Sinha, Arunesh
Varakantham, Pradeep
author_facet Lu, Yuxiao
Sinha, Arunesh
Varakantham, Pradeep
contents Safety in goal directed Reinforcement Learning (RL) settings has typically been handled through constraints over trajectories and have demonstrated good performance in primarily short horizon tasks. In this paper, we are specifically interested in the problem of solving temporally extended decision making problems such as robots cleaning different areas in a house while avoiding slippery and unsafe areas (e.g., stairs) and retaining enough charge to move to a charging dock; in the presence of complex safety constraints. Our key contribution is a (safety) Constrained Search with Hierarchical Reinforcement Learning (CoSHRL) mechanism that combines an upper level constrained search agent (which computes a reward maximizing policy from a given start to a far away goal state while satisfying cost constraints) with a low-level goal conditioned RL agent (which estimates cost and reward values to move between nearby states). A major advantage of CoSHRL is that it can handle constraints on the cost value distribution (e.g., on Conditional Value at Risk, CVaR) and can adjust to flexible constraint thresholds without retraining. We perform extensive experiments with different types of safety constraints to demonstrate the utility of our approach over leading approaches in constrained and hierarchical RL.
format Preprint
id arxiv_https___arxiv_org_abs_2302_10639
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
Lu, Yuxiao
Sinha, Arunesh
Varakantham, Pradeep
Artificial Intelligence
Machine Learning
Safety in goal directed Reinforcement Learning (RL) settings has typically been handled through constraints over trajectories and have demonstrated good performance in primarily short horizon tasks. In this paper, we are specifically interested in the problem of solving temporally extended decision making problems such as robots cleaning different areas in a house while avoiding slippery and unsafe areas (e.g., stairs) and retaining enough charge to move to a charging dock; in the presence of complex safety constraints. Our key contribution is a (safety) Constrained Search with Hierarchical Reinforcement Learning (CoSHRL) mechanism that combines an upper level constrained search agent (which computes a reward maximizing policy from a given start to a far away goal state while satisfying cost constraints) with a low-level goal conditioned RL agent (which estimates cost and reward values to move between nearby states). A major advantage of CoSHRL is that it can handle constraints on the cost value distribution (e.g., on Conditional Value at Risk, CVaR) and can adjust to flexible constraint thresholds without retraining. We perform extensive experiments with different types of safety constraints to demonstrate the utility of our approach over leading approaches in constrained and hierarchical RL.
title Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2302.10639