EgoMimic: Scaling Imitation Learning via Egocentric Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kareer, Simar, Patel, Dhruv, Punamiya, Ryan, Mathur, Pranay, Cheng, Shuo, Wang, Chen, Hoffman, Judy, Xu, Danfei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915001160171520
author Kareer, Simar
Patel, Dhruv
Punamiya, Ryan
Mathur, Pranay
Cheng, Shuo
Wang, Chen
Hoffman, Judy
Xu, Danfei
author_facet Kareer, Simar
Patel, Dhruv
Punamiya, Ryan
Mathur, Pranay
Cheng, Shuo
Wang, Chen
Hoffman, Judy
Xu, Danfei
contents The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2410_24221
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EgoMimic: Scaling Imitation Learning via Egocentric Video
Kareer, Simar
Patel, Dhruv
Punamiya, Ryan
Mathur, Pranay
Cheng, Shuo
Wang, Chen
Hoffman, Judy
Xu, Danfei
Robotics
Computer Vision and Pattern Recognition
The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/
title EgoMimic: Scaling Imitation Learning via Egocentric Video
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.24221