DITTO: Demonstration Imitation by Trajectory Transformation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heppert, Nick, Argus, Max, Welschehold, Tim, Brox, Thomas, Valada, Abhinav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910803878215680
author Heppert, Nick
Argus, Max
Welschehold, Tim
Brox, Thomas
Valada, Abhinav
author_facet Heppert, Nick
Argus, Max
Welschehold, Tim
Brox, Thomas
Valada, Abhinav
contents Teaching robots new skills quickly and conveniently is crucial for the broader adoption of robotic systems. In this work, we address the problem of one-shot imitation from a single human demonstration, given by an RGB-D video recording. We propose a two-stage process. In the first stage we extract the demonstration trajectory offline. This entails segmenting manipulated objects and determining their relative motion in relation to secondary objects such as containers. In the online trajectory generation stage, we first re-detect all objects, then warp the demonstration trajectory to the current scene and execute it on the robot. To complete these steps, our method leverages several ancillary models, including those for segmentation, relative object pose estimation, and grasp prediction. We systematically evaluate different combinations of correspondence and re-detection methods to validate our design decision across a diverse range of tasks. Specifically, we collect and quantitatively test on demonstrations of ten different tasks including pick-and-place tasks as well as articulated object manipulation. Finally, we perform extensive evaluations on a real robot system to demonstrate the effectiveness and utility of our approach in real-world scenarios. We make the code publicly available at http://ditto.cs.uni-freiburg.de.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DITTO: Demonstration Imitation by Trajectory Transformation
Heppert, Nick
Argus, Max
Welschehold, Tim
Brox, Thomas
Valada, Abhinav
Robotics
Computer Vision and Pattern Recognition
Teaching robots new skills quickly and conveniently is crucial for the broader adoption of robotic systems. In this work, we address the problem of one-shot imitation from a single human demonstration, given by an RGB-D video recording. We propose a two-stage process. In the first stage we extract the demonstration trajectory offline. This entails segmenting manipulated objects and determining their relative motion in relation to secondary objects such as containers. In the online trajectory generation stage, we first re-detect all objects, then warp the demonstration trajectory to the current scene and execute it on the robot. To complete these steps, our method leverages several ancillary models, including those for segmentation, relative object pose estimation, and grasp prediction. We systematically evaluate different combinations of correspondence and re-detection methods to validate our design decision across a diverse range of tasks. Specifically, we collect and quantitatively test on demonstrations of ten different tasks including pick-and-place tasks as well as articulated object manipulation. Finally, we perform extensive evaluations on a real robot system to demonstrate the effectiveness and utility of our approach in real-world scenarios. We make the code publicly available at http://ditto.cs.uni-freiburg.de.
title DITTO: Demonstration Imitation by Trajectory Transformation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.15203