Salvato in:
Dettagli Bibliografici
Autori principali: Li, Jinhan, Zhu, Yifeng, Xie, Yuqi, Jiang, Zhenyu, Seo, Mingyo, Pavlakos, Georgios, Zhu, Yuke
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2410.11792
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929544076722176
author Li, Jinhan
Zhu, Yifeng
Xie, Yuqi
Jiang, Zhenyu
Seo, Mingyo
Pavlakos, Georgios
Zhu, Yuke
author_facet Li, Jinhan
Zhu, Yifeng
Xie, Yuqi
Jiang, Zhenyu
Seo, Mingyo
Pavlakos, Georgios
Zhu, Yuke
contents We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting, which enables the humanoid robot to mimic the human motions in an RGB-D video while adjusting to different object locations during deployment. OKAMI uses open-world vision models to identify task-relevant objects and retarget the body motions and hand poses separately. Our experiments show that OKAMI achieves strong generalizations across varying visual and spatial conditions, outperforming the state-of-the-art baseline on open-world imitation from observation. Furthermore, OKAMI rollout trajectories are leveraged to train closed-loop visuomotor policies, which achieve an average success rate of 79.2% without the need for labor-intensive teleoperation. More videos can be found on our website https://ut-austin-rpl.github.io/OKAMI/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11792
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
Li, Jinhan
Zhu, Yifeng
Xie, Yuqi
Jiang, Zhenyu
Seo, Mingyo
Pavlakos, Georgios
Zhu, Yuke
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting, which enables the humanoid robot to mimic the human motions in an RGB-D video while adjusting to different object locations during deployment. OKAMI uses open-world vision models to identify task-relevant objects and retarget the body motions and hand poses separately. Our experiments show that OKAMI achieves strong generalizations across varying visual and spatial conditions, outperforming the state-of-the-art baseline on open-world imitation from observation. Furthermore, OKAMI rollout trajectories are leveraged to train closed-loop visuomotor policies, which achieve an average success rate of 79.2% without the need for labor-intensive teleoperation. More videos can be found on our website https://ut-austin-rpl.github.io/OKAMI/.
title OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.11792