EgoMimic: Scaling Imitation Learning via Egocentric Video
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915001160171520 |
|---|---|
| author | Kareer, Simar Patel, Dhruv Punamiya, Ryan Mathur, Pranay Cheng, Shuo Wang, Chen Hoffman, Judy Xu, Danfei |
| author_facet | Kareer, Simar Patel, Dhruv Punamiya, Ryan Mathur, Pranay Cheng, Shuo Wang, Chen Hoffman, Judy Xu, Danfei |
| contents | The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_24221 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | EgoMimic: Scaling Imitation Learning via Egocentric Video Kareer, Simar Patel, Dhruv Punamiya, Ryan Mathur, Pranay Cheng, Shuo Wang, Chen Hoffman, Judy Xu, Danfei Robotics Computer Vision and Pattern Recognition The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/ |
| title | EgoMimic: Scaling Imitation Learning via Egocentric Video |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.24221 |