A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Qihan, Yang, Xinsong, Xia, Gang, Ho, Daniel W. C., Tang, Pengyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916429295517696
author Qi, Qihan
Yang, Xinsong
Xia, Gang
Ho, Daniel W. C.
Tang, Pengyang
author_facet Qi, Qihan
Yang, Xinsong
Xia, Gang
Ho, Daniel W. C.
Tang, Pengyang
contents This paper proposes a safety modulator actor-critic (SMAC) method to address safety constraint and overestimation mitigation in model-free safe reinforcement learning (RL). A safety modulator is developed to satisfy safety constraints by modulating actions, allowing the policy to ignore safety constraint and focus on maximizing reward. Additionally, a distributional critic with a theoretical update rule for SMAC is proposed to mitigate the overestimation of Q-values with safety constraints. Both simulation and real-world scenarios experiments on Unmanned Aerial Vehicles (UAVs) hovering confirm that the SMAC can effectively maintain safety constraints and outperform mainstream baseline algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06847
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering
Qi, Qihan
Yang, Xinsong
Xia, Gang
Ho, Daniel W. C.
Tang, Pengyang
Artificial Intelligence
Machine Learning
Robotics
This paper proposes a safety modulator actor-critic (SMAC) method to address safety constraint and overestimation mitigation in model-free safe reinforcement learning (RL). A safety modulator is developed to satisfy safety constraints by modulating actions, allowing the policy to ignore safety constraint and focus on maximizing reward. Additionally, a distributional critic with a theoretical update rule for SMAC is proposed to mitigate the overestimation of Q-values with safety constraints. Both simulation and real-world scenarios experiments on Unmanned Aerial Vehicles (UAVs) hovering confirm that the SMAC can effectively maintain safety constraints and outperform mainstream baseline algorithms.
title A Safety Modulator Actor-Critic Method in Model-Free Safe Reinforcement Learning and Application in UAV Hovering
topic Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2410.06847