Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Puhao, Liu, Tengyu, Li, Yuyang, Han, Muzhi, Geng, Haoran, Wang, Shu, Zhu, Yixin, Zhu, Song-Chun, Huang, Siyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866907882975395840
author Li, Puhao
Liu, Tengyu
Li, Yuyang
Han, Muzhi
Geng, Haoran
Wang, Shu
Zhu, Yixin
Zhu, Song-Chun
Huang, Siyuan
author_facet Li, Puhao
Liu, Tengyu
Li, Yuyang
Han, Muzhi
Geng, Haoran
Wang, Shu
Zhu, Yixin
Zhu, Song-Chun
Huang, Siyuan
contents Autonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, modern methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of successful task executions within specific action spaces, resulting in misaligned and ambiguous task representations. We introduce Ag2Manip (Agent-Agnostic representations for Manipulation), a framework aimed at surmounting these challenges through two key innovations: a novel agent-agnostic visual representation derived from human manipulation videos, with the specifics of embodiments obscured to enhance generalizability; and an agent-agnostic action representation abstracting a robot's kinematics to a universal agent proxy, emphasizing crucial interactions between end-effector and object. Ag2Manip's empirical validation across simulated benchmarks like FrankaKitchen, ManiSkill, and PartManip shows a 325% increase in performance, achieved without domain-specific demonstrations. Ablation studies underline the essential contributions of the visual and action representations to this success. Extending our evaluations to the real world, Ag2Manip significantly improves imitation learning success rates from 50% to 77.5%, demonstrating its effectiveness and generalizability across both simulated and physical environments.
format Preprint
id arxiv_https___arxiv_org_abs_2404_17521
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
Li, Puhao
Liu, Tengyu
Li, Yuyang
Han, Muzhi
Geng, Haoran
Wang, Shu
Zhu, Yixin
Zhu, Song-Chun
Huang, Siyuan
Robotics
Computer Vision and Pattern Recognition
Autonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, modern methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of successful task executions within specific action spaces, resulting in misaligned and ambiguous task representations. We introduce Ag2Manip (Agent-Agnostic representations for Manipulation), a framework aimed at surmounting these challenges through two key innovations: a novel agent-agnostic visual representation derived from human manipulation videos, with the specifics of embodiments obscured to enhance generalizability; and an agent-agnostic action representation abstracting a robot's kinematics to a universal agent proxy, emphasizing crucial interactions between end-effector and object. Ag2Manip's empirical validation across simulated benchmarks like FrankaKitchen, ManiSkill, and PartManip shows a 325% increase in performance, achieved without domain-specific demonstrations. Ablation studies underline the essential contributions of the visual and action representations to this success. Extending our evaluations to the real world, Ag2Manip significantly improves imitation learning success rates from 50% to 77.5%, demonstrating its effectiveness and generalizability across both simulated and physical environments.
title Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.17521