Reinforcement Learning from Wild Animal Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chane-Sane, Elliot, Roux, Constant, Stasse, Olivier, Mansard, Nicolas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916509219028992
author Chane-Sane, Elliot
Roux, Constant
Stasse, Olivier
Mansard, Nicolas
author_facet Chane-Sane, Elliot
Roux, Constant
Stasse, Olivier
Mansard, Nicolas
contents We propose to learn legged robot locomotion skills by watching thousands of wild animal videos from the internet, such as those featured in nature documentaries. Indeed, such videos offer a rich and diverse collection of plausible motion examples, which could inform how robots should move. To achieve this, we introduce Reinforcement Learning from Wild Animal Videos (RLWAV), a method to ground these motions into physical robots. We first train a video classifier on a large-scale animal video dataset to recognize actions from RGB clips of animals in their natural habitats. We then train a multi-skill policy to control a robot in a physics simulator, using the classification score of a third-person camera capturing videos of the robot's movements as a reward for reinforcement learning. Finally, we directly transfer the learned policy to a real quadruped Solo. Remarkably, despite the extreme gap in both domain and embodiment between animals in the wild and robots, our approach enables the policy to learn diverse skills such as walking, jumping, and keeping still, without relying on reference trajectories nor skill-specific rewards.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04273
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reinforcement Learning from Wild Animal Videos
Chane-Sane, Elliot
Roux, Constant
Stasse, Olivier
Mansard, Nicolas
Robotics
Computer Vision and Pattern Recognition
Machine Learning
We propose to learn legged robot locomotion skills by watching thousands of wild animal videos from the internet, such as those featured in nature documentaries. Indeed, such videos offer a rich and diverse collection of plausible motion examples, which could inform how robots should move. To achieve this, we introduce Reinforcement Learning from Wild Animal Videos (RLWAV), a method to ground these motions into physical robots. We first train a video classifier on a large-scale animal video dataset to recognize actions from RGB clips of animals in their natural habitats. We then train a multi-skill policy to control a robot in a physics simulator, using the classification score of a third-person camera capturing videos of the robot's movements as a reward for reinforcement learning. Finally, we directly transfer the learned policy to a real quadruped Solo. Remarkably, despite the extreme gap in both domain and embodiment between animals in the wild and robots, our approach enables the policy to learn diverse skills such as walking, jumping, and keeping still, without relying on reference trajectories nor skill-specific rewards.
title Reinforcement Learning from Wild Animal Videos
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.04273