Visual Imitation Learning of Task-Oriented Object Grasping and Rearrangement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Yichen, Gao, Jianfeng, Pohl, Christoph, Asfour, Tamim
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908866233499648
author Cai, Yichen
Gao, Jianfeng
Pohl, Christoph
Asfour, Tamim
author_facet Cai, Yichen
Gao, Jianfeng
Pohl, Christoph
Asfour, Tamim
contents Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in categorical objects. In this paper, we propose the Multi-feature Implicit Model (MIMO), a novel object representation that encodes multiple spatial features between a point and an object in an implicit neural field. Training such a model on multiple features ensures that it embeds the object shapes consistently in different aspects, thus improving its performance in object shape reconstruction from partial observation, shape similarity measure, and modeling spatial relations between objects. Based on MIMO, we propose a framework to learn task-oriented object grasping and rearrangement from single or multiple human demonstration videos. The evaluations in simulation show that our approach outperforms the state-of-the-art methods for multi- and single-view observations. Real-world experiments demonstrate the efficacy of our approach in one- and few-shot imitation learning of manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14000
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visual Imitation Learning of Task-Oriented Object Grasping and Rearrangement
Cai, Yichen
Gao, Jianfeng
Pohl, Christoph
Asfour, Tamim
Robotics
Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in categorical objects. In this paper, we propose the Multi-feature Implicit Model (MIMO), a novel object representation that encodes multiple spatial features between a point and an object in an implicit neural field. Training such a model on multiple features ensures that it embeds the object shapes consistently in different aspects, thus improving its performance in object shape reconstruction from partial observation, shape similarity measure, and modeling spatial relations between objects. Based on MIMO, we propose a framework to learn task-oriented object grasping and rearrangement from single or multiple human demonstration videos. The evaluations in simulation show that our approach outperforms the state-of-the-art methods for multi- and single-view observations. Real-world experiments demonstrate the efficacy of our approach in one- and few-shot imitation learning of manipulation tasks.
title Visual Imitation Learning of Task-Oriented Object Grasping and Rearrangement
topic Robotics
url https://arxiv.org/abs/2403.14000