One Policy to Run Them All: an End-to-end Learning Approach to Multi-Embodiment Locomotion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bohlinger, Nico, Czechmanowski, Grzegorz, Krupka, Maciej, Kicki, Piotr, Walas, Krzysztof, Peters, Jan, Tateo, Davide
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914070902341632
author Bohlinger, Nico
Czechmanowski, Grzegorz
Krupka, Maciej
Kicki, Piotr
Walas, Krzysztof
Peters, Jan
Tateo, Davide
author_facet Bohlinger, Nico
Czechmanowski, Grzegorz
Krupka, Maciej
Kicki, Piotr
Walas, Krzysztof
Peters, Jan
Tateo, Davide
contents Deep Reinforcement Learning techniques are achieving state-of-the-art results in robust legged locomotion. While there exists a wide variety of legged platforms such as quadruped, humanoids, and hexapods, the field is still missing a single learning framework that can control all these different embodiments easily and effectively and possibly transfer, zero or few-shot, to unseen robot embodiments. We introduce URMA, the Unified Robot Morphology Architecture, to close this gap. Our framework brings the end-to-end Multi-Task Reinforcement Learning approach to the realm of legged robots, enabling the learned policy to control any type of robot morphology. The key idea of our method is to allow the network to learn an abstract locomotion controller that can be seamlessly shared between embodiments thanks to our morphology-agnostic encoders and decoders. This flexible architecture can be seen as a potential first step in building a foundation model for legged robot locomotion. Our experiments show that URMA can learn a locomotion policy on multiple embodiments that can be easily transferred to unseen robot platforms in simulation and the real world.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle One Policy to Run Them All: an End-to-end Learning Approach to Multi-Embodiment Locomotion
Bohlinger, Nico
Czechmanowski, Grzegorz
Krupka, Maciej
Kicki, Piotr
Walas, Krzysztof
Peters, Jan
Tateo, Davide
Robotics
Machine Learning
Deep Reinforcement Learning techniques are achieving state-of-the-art results in robust legged locomotion. While there exists a wide variety of legged platforms such as quadruped, humanoids, and hexapods, the field is still missing a single learning framework that can control all these different embodiments easily and effectively and possibly transfer, zero or few-shot, to unseen robot embodiments. We introduce URMA, the Unified Robot Morphology Architecture, to close this gap. Our framework brings the end-to-end Multi-Task Reinforcement Learning approach to the realm of legged robots, enabling the learned policy to control any type of robot morphology. The key idea of our method is to allow the network to learn an abstract locomotion controller that can be seamlessly shared between embodiments thanks to our morphology-agnostic encoders and decoders. This flexible architecture can be seen as a potential first step in building a foundation model for legged robot locomotion. Our experiments show that URMA can learn a locomotion policy on multiple embodiments that can be easily transferred to unseen robot platforms in simulation and the real world.
title One Policy to Run Them All: an End-to-end Learning Approach to Multi-Embodiment Locomotion
topic Robotics
Machine Learning
url https://arxiv.org/abs/2409.06366