Benchmarking Population-Based Reinforcement Learning across Robotic Tasks with GPU-Accelerated Simulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shahid, Asad Ali, Narang, Yashraj, Petrone, Vincenzo, Ferrentino, Enrico, Handa, Ankur, Fox, Dieter, Pavone, Marco, Roveda, Loris
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911272082079744
author Shahid, Asad Ali
Narang, Yashraj
Petrone, Vincenzo
Ferrentino, Enrico
Handa, Ankur
Fox, Dieter
Pavone, Marco
Roveda, Loris
author_facet Shahid, Asad Ali
Narang, Yashraj
Petrone, Vincenzo
Ferrentino, Enrico
Handa, Ankur
Fox, Dieter
Pavone, Marco
Roveda, Loris
contents In recent years, deep reinforcement learning (RL) has shown its effectiveness in solving complex continuous control tasks. However, this comes at the cost of an enormous amount of experience required for training, exacerbated by the sensitivity of learning efficiency and the policy performance to hyperparameter selection, which often requires numerous trials of time-consuming experiments. This work leverages a Population-Based Reinforcement Learning (PBRL) approach and a GPU-accelerated physics simulator to enhance the exploration capabilities of RL by concurrently training multiple policies in parallel. The PBRL framework is benchmarked against three state-of-the-art RL algorithms -- PPO, SAC, and DDPG -- dynamically adjusting hyperparameters based on the performance of learning agents. The experiments are performed on four challenging tasks in Isaac Gym -- Anymal Terrain, Shadow Hand, Humanoid, Franka Nut Pick -- by analyzing the effect of population size and mutation mechanisms for hyperparameters. The results show that PBRL agents achieve superior performance, in terms of cumulative reward, compared to non-evolutionary baseline agents. Moreover, the trained agents are finally deployed in the real world for a Franka Nut Pick task. To our knowledge, this is the first sim-to-real attempt for deploying PBRL agents on real hardware. Code and videos of the learned policies are available on our project website (https://sites.google.com/view/pbrl).
format Preprint
id arxiv_https___arxiv_org_abs_2404_03336
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Population-Based Reinforcement Learning across Robotic Tasks with GPU-Accelerated Simulation
Shahid, Asad Ali
Narang, Yashraj
Petrone, Vincenzo
Ferrentino, Enrico
Handa, Ankur
Fox, Dieter
Pavone, Marco
Roveda, Loris
Robotics
In recent years, deep reinforcement learning (RL) has shown its effectiveness in solving complex continuous control tasks. However, this comes at the cost of an enormous amount of experience required for training, exacerbated by the sensitivity of learning efficiency and the policy performance to hyperparameter selection, which often requires numerous trials of time-consuming experiments. This work leverages a Population-Based Reinforcement Learning (PBRL) approach and a GPU-accelerated physics simulator to enhance the exploration capabilities of RL by concurrently training multiple policies in parallel. The PBRL framework is benchmarked against three state-of-the-art RL algorithms -- PPO, SAC, and DDPG -- dynamically adjusting hyperparameters based on the performance of learning agents. The experiments are performed on four challenging tasks in Isaac Gym -- Anymal Terrain, Shadow Hand, Humanoid, Franka Nut Pick -- by analyzing the effect of population size and mutation mechanisms for hyperparameters. The results show that PBRL agents achieve superior performance, in terms of cumulative reward, compared to non-evolutionary baseline agents. Moreover, the trained agents are finally deployed in the real world for a Franka Nut Pick task. To our knowledge, this is the first sim-to-real attempt for deploying PBRL agents on real hardware. Code and videos of the learned policies are available on our project website (https://sites.google.com/view/pbrl).
title Benchmarking Population-Based Reinforcement Learning across Robotic Tasks with GPU-Accelerated Simulation
topic Robotics
url https://arxiv.org/abs/2404.03336