SDS -- See it, Do it, Sorted: Quadruped Skill Synthesis from Single Video Demonstration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stamatopoulou, Maria, Li, Jeffrey, Kanoulas, Dimitrios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916909561151488
author Stamatopoulou, Maria
Li, Jeffrey
Kanoulas, Dimitrios
author_facet Stamatopoulou, Maria
Li, Jeffrey
Kanoulas, Dimitrios
contents Imagine a robot learning locomotion skills from any single video, without labels or reward engineering. We introduce SDS ("See it. Do it. Sorted."), an automated pipeline for skill acquisition from unstructured demonstrations. Using GPT-4o, SDS applies novel prompting techniques, in the form of spatio-temporal grid-based visual encoding ($G_{v}$) and structured input decomposition (SUS). These produce executable reward functions (RF) from the raw input videos. The RFs are used to train PPO policies and are optimized through closed-loop evolution, using training footage and performance metrics as self-supervised signals. SDS allows quadrupeds (e.g. Unitree Go1) to learn four gaits -- trot, bound, pace, and hop -- achieving 100% gait matching fidelity, Dynamic Time Warping (DTW) distance in the order of $10^{-6}$, and stable locomotion with zero failures, both in simulation and the real world. SDS generalizes to morphologically different quadrupeds (e.g. ANYmal) and outperforms prior work in data efficiency, training time and engineering effort. Further materials and the code are open-source under: https://rpl-cs-ucl.github.io/SDSweb/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11571
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SDS -- See it, Do it, Sorted: Quadruped Skill Synthesis from Single Video Demonstration
Stamatopoulou, Maria
Li, Jeffrey
Kanoulas, Dimitrios
Robotics
Imagine a robot learning locomotion skills from any single video, without labels or reward engineering. We introduce SDS ("See it. Do it. Sorted."), an automated pipeline for skill acquisition from unstructured demonstrations. Using GPT-4o, SDS applies novel prompting techniques, in the form of spatio-temporal grid-based visual encoding ($G_{v}$) and structured input decomposition (SUS). These produce executable reward functions (RF) from the raw input videos. The RFs are used to train PPO policies and are optimized through closed-loop evolution, using training footage and performance metrics as self-supervised signals. SDS allows quadrupeds (e.g. Unitree Go1) to learn four gaits -- trot, bound, pace, and hop -- achieving 100% gait matching fidelity, Dynamic Time Warping (DTW) distance in the order of $10^{-6}$, and stable locomotion with zero failures, both in simulation and the real world. SDS generalizes to morphologically different quadrupeds (e.g. ANYmal) and outperforms prior work in data efficiency, training time and engineering effort. Further materials and the code are open-source under: https://rpl-cs-ucl.github.io/SDSweb/.
title SDS -- See it, Do it, Sorted: Quadruped Skill Synthesis from Single Video Demonstration
topic Robotics
url https://arxiv.org/abs/2410.11571