Safe Reinforcement Learning in a Simulated Robotic Arm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kovač, Luka, Farkaš, Igor
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917600430129152
author Kovač, Luka
Farkaš, Igor
author_facet Kovač, Luka
Farkaš, Igor
contents Reinforcement learning (RL) agents need to explore their environments in order to learn optimal policies. In many environments and tasks, safety is of critical importance. The widespread use of simulators offers a number of advantages, including safe exploration which will be inevitable in cases when RL systems need to be trained directly in the physical environment (e.g. in human-robot interaction). The popular Safety Gym library offers three mobile agent types that can learn goal-directed tasks while considering various safety constraints. In this paper, we extend the applicability of safe RL algorithms by creating a customized environment with Panda robotic arm where Safety Gym algorithms can be tested. We performed pilot experiments with the popular PPO algorithm comparing the baseline with the constrained version and show that the constrained version is able to learn the equally good policy while better complying with safety constraints and taking longer training time as expected.
format Preprint
id arxiv_https___arxiv_org_abs_2312_09468
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Safe Reinforcement Learning in a Simulated Robotic Arm
Kovač, Luka
Farkaš, Igor
Robotics
Artificial Intelligence
Machine Learning
Reinforcement learning (RL) agents need to explore their environments in order to learn optimal policies. In many environments and tasks, safety is of critical importance. The widespread use of simulators offers a number of advantages, including safe exploration which will be inevitable in cases when RL systems need to be trained directly in the physical environment (e.g. in human-robot interaction). The popular Safety Gym library offers three mobile agent types that can learn goal-directed tasks while considering various safety constraints. In this paper, we extend the applicability of safe RL algorithms by creating a customized environment with Panda robotic arm where Safety Gym algorithms can be tested. We performed pilot experiments with the popular PPO algorithm comparing the baseline with the constrained version and show that the constrained version is able to learn the equally good policy while better complying with safety constraints and taking longer training time as expected.
title Safe Reinforcement Learning in a Simulated Robotic Arm
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.09468