Vision-based Manipulation from Single Human Video with Open-World Object Graphs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yifeng, Lim, Arisrei, Stone, Peter, Zhu, Yuke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914020773068800
author Zhu, Yifeng
Lim, Arisrei
Stone, Peter
Zhu, Yuke
author_facet Zhu, Yifeng
Lim, Arisrei
Stone, Peter
Zhu, Yuke
contents This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel objects from a single video demonstration. We introduce ORION, an algorithm that tackles the problem by extracting an object-centric manipulation plan from a single RGB or RGB-D video and deriving a policy that conditions on the extracted plan. Our method enables the robot to learn from videos captured by daily mobile devices and to generalize the policies to deployment environments with varying visual backgrounds, camera angles, spatial layouts, and novel object instances. We systematically evaluate our method on both short-horizon and long-horizon tasks, using RGB-D and RGB-only demonstration videos. Across varied tasks and demonstration types (RGB-D / RGB), we observe an average success rate of 74.4%, demonstrating the efficacy of ORION in learning from a single human video in the open world. Additional materials can be found on our project website: https://ut-austin-rpl.github.io/ORION-release.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vision-based Manipulation from Single Human Video with Open-World Object Graphs
Zhu, Yifeng
Lim, Arisrei
Stone, Peter
Zhu, Yuke
Robotics
Computer Vision and Pattern Recognition
Machine Learning
This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel objects from a single video demonstration. We introduce ORION, an algorithm that tackles the problem by extracting an object-centric manipulation plan from a single RGB or RGB-D video and deriving a policy that conditions on the extracted plan. Our method enables the robot to learn from videos captured by daily mobile devices and to generalize the policies to deployment environments with varying visual backgrounds, camera angles, spatial layouts, and novel object instances. We systematically evaluate our method on both short-horizon and long-horizon tasks, using RGB-D and RGB-only demonstration videos. Across varied tasks and demonstration types (RGB-D / RGB), we observe an average success rate of 74.4%, demonstrating the efficacy of ORION in learning from a single human video in the open world. Additional materials can be found on our project website: https://ut-austin-rpl.github.io/ORION-release.
title Vision-based Manipulation from Single Human Video with Open-World Object Graphs
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.20321