An Open-Loop Baseline for Reinforcement Learning Locomotion Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raffin, Antonin, Sigaud, Olivier, Kober, Jens, Albu-Schäffer, Alin, Silvério, João, Stulp, Freek
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914700395020288
author Raffin, Antonin
Sigaud, Olivier
Kober, Jens
Albu-Schäffer, Alin
Silvério, João
Stulp, Freek
author_facet Raffin, Antonin
Sigaud, Olivier
Kober, Jens
Albu-Schäffer, Alin
Silvério, João
Stulp, Freek
contents In search of a simple baseline for Deep Reinforcement Learning in locomotion tasks, we propose a model-free open-loop strategy. By leveraging prior knowledge and the elegance of simple oscillators to generate periodic joint motions, it achieves respectable performance in five different locomotion environments, with a number of tunable parameters that is a tiny fraction of the thousands typically required by DRL algorithms. We conduct two additional experiments using open-loop oscillators to identify current shortcomings of these algorithms. Our results show that, compared to the baseline, DRL is more prone to performance degradation when exposed to sensor noise or failure. Furthermore, we demonstrate a successful transfer from simulation to reality using an elastic quadruped, where RL fails without randomization or reward engineering. Overall, the proposed baseline and associated experiments highlight the existing limitations of DRL for robotic applications, provide insights on how to address them, and encourage reflection on the costs of complexity and generality.
format Preprint
id arxiv_https___arxiv_org_abs_2310_05808
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle An Open-Loop Baseline for Reinforcement Learning Locomotion Tasks
Raffin, Antonin
Sigaud, Olivier
Kober, Jens
Albu-Schäffer, Alin
Silvério, João
Stulp, Freek
Robotics
In search of a simple baseline for Deep Reinforcement Learning in locomotion tasks, we propose a model-free open-loop strategy. By leveraging prior knowledge and the elegance of simple oscillators to generate periodic joint motions, it achieves respectable performance in five different locomotion environments, with a number of tunable parameters that is a tiny fraction of the thousands typically required by DRL algorithms. We conduct two additional experiments using open-loop oscillators to identify current shortcomings of these algorithms. Our results show that, compared to the baseline, DRL is more prone to performance degradation when exposed to sensor noise or failure. Furthermore, we demonstrate a successful transfer from simulation to reality using an elastic quadruped, where RL fails without randomization or reward engineering. Overall, the proposed baseline and associated experiments highlight the existing limitations of DRL for robotic applications, provide insights on how to address them, and encourage reflection on the costs of complexity and generality.
title An Open-Loop Baseline for Reinforcement Learning Locomotion Tasks
topic Robotics
url https://arxiv.org/abs/2310.05808