Guaranteeing Control Requirements via Reward Shaping in Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De Lellis, Francesco, Coraggio, Marco, Russo, Giovanni, Musolesi, Mirco, di Bernardo, Mario
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917618508627968
author De Lellis, Francesco
Coraggio, Marco
Russo, Giovanni
Musolesi, Mirco
di Bernardo, Mario
author_facet De Lellis, Francesco
Coraggio, Marco
Russo, Giovanni
Musolesi, Mirco
di Bernardo, Mario
contents In addressing control problems such as regulation and tracking through reinforcement learning, it is often required to guarantee that the acquired policy meets essential performance and stability criteria such as a desired settling time and steady-state error prior to deployment. Motivated by this necessity, we present a set of results and a systematic reward shaping procedure that (i) ensures the optimal policy generates trajectories that align with specified control requirements and (ii) allows to assess whether any given policy satisfies them. We validate our approach through comprehensive numerical experiments conducted in two representative environments from OpenAI Gym: the Inverted Pendulum swing-up problem and the Lunar Lander. Utilizing both tabular and deep reinforcement learning methods, our experiments consistently affirm the efficacy of our proposed framework, highlighting its effectiveness in ensuring policy adherence to the prescribed control requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2311_10026
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Guaranteeing Control Requirements via Reward Shaping in Reinforcement Learning
De Lellis, Francesco
Coraggio, Marco
Russo, Giovanni
Musolesi, Mirco
di Bernardo, Mario
Systems and Control
Machine Learning
In addressing control problems such as regulation and tracking through reinforcement learning, it is often required to guarantee that the acquired policy meets essential performance and stability criteria such as a desired settling time and steady-state error prior to deployment. Motivated by this necessity, we present a set of results and a systematic reward shaping procedure that (i) ensures the optimal policy generates trajectories that align with specified control requirements and (ii) allows to assess whether any given policy satisfies them. We validate our approach through comprehensive numerical experiments conducted in two representative environments from OpenAI Gym: the Inverted Pendulum swing-up problem and the Lunar Lander. Utilizing both tabular and deep reinforcement learning methods, our experiments consistently affirm the efficacy of our proposed framework, highlighting its effectiveness in ensuring policy adherence to the prescribed control requirements.
title Guaranteeing Control Requirements via Reward Shaping in Reinforcement Learning
topic Systems and Control
Machine Learning
url https://arxiv.org/abs/2311.10026