VIHE: Virtual In-Hand Eye Transformer for 3D Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Weiyao, Lei, Yutian, Jin, Shiyu, Hager, Gregory D., Zhang, Liangjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913271470096384
author Wang, Weiyao
Lei, Yutian
Jin, Shiyu
Hager, Gregory D.
Zhang, Liangjun
author_facet Wang, Weiyao
Lei, Yutian
Jin, Shiyu
Hager, Gregory D.
Zhang, Liangjun
contents In this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong inductive bias for effectively recognizing the correct pose for the hand, especially for challenging high-precision tasks such as peg insertion. On 18 manipulation tasks in RLBench simulated environments, VIHE achieves a new state-of-the-art, with a 12% absolute improvement, increasing from 65% to 77% over the existing state-of-the-art model using 100 demonstrations per task. In real-world scenarios, VIHE can learn manipulation tasks with just a handful of demonstrations, highlighting its practical utility. Videos and code implementation can be found at our project site: https://vihe-3d.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2403_11461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VIHE: Virtual In-Hand Eye Transformer for 3D Robotic Manipulation
Wang, Weiyao
Lei, Yutian
Jin, Shiyu
Hager, Gregory D.
Zhang, Liangjun
Robotics
In this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong inductive bias for effectively recognizing the correct pose for the hand, especially for challenging high-precision tasks such as peg insertion. On 18 manipulation tasks in RLBench simulated environments, VIHE achieves a new state-of-the-art, with a 12% absolute improvement, increasing from 65% to 77% over the existing state-of-the-art model using 100 demonstrations per task. In real-world scenarios, VIHE can learn manipulation tasks with just a handful of demonstrations, highlighting its practical utility. Videos and code implementation can be found at our project site: https://vihe-3d.github.io.
title VIHE: Virtual In-Hand Eye Transformer for 3D Robotic Manipulation
topic Robotics
url https://arxiv.org/abs/2403.11461