Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yang, Yang, Sizhe, Zeng, Jia, Wang, Ping, Lin, Dahua, Dong, Hao, Pang, Jiangmiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929641383526400
author Tian, Yang
Yang, Sizhe
Zeng, Jia
Wang, Ping
Lin, Dahua
Dong, Hao
Pang, Jiangmiao
author_facet Tian, Yang
Yang, Sizhe
Zeng, Jia
Wang, Ping
Lin, Dahua
Dong, Hao
Pang, Jiangmiao
contents Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representations or generative models, also referred to as world models, using large-scale visual datasets. This paper presents an end-to-end paradigm that predicts actions using inverse dynamics models conditioned on the robot's forecasted visual states, named Predictive Inverse Dynamics Models (PIDM). By closing the loop between vision and action, the end-to-end PIDM can be a better scalable action learner. In practice, we use Transformers to process both visual states and actions, naming the model Seer. It is initially pre-trained on large-scale robotic datasets, such as DROID, and can be adapted to realworld scenarios with a little fine-tuning data. Thanks to large-scale, end-to-end training and the synergy between vision and action, Seer significantly outperforms previous methods across both simulation and real-world experiments. It achieves improvements of 13% on the LIBERO-LONG benchmark, 21% on CALVIN ABC-D, and 43% in real-world tasks. Notably, Seer sets a new state-of-the-art on CALVIN ABC-D benchmark, achieving an average length of 4.28, and exhibits superior generalization for novel objects, lighting conditions, and environments under high-intensity disturbances on real-world scenarios. Code and models are publicly available at https://github.com/OpenRobotLab/Seer/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15109
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
Tian, Yang
Yang, Sizhe
Zeng, Jia
Wang, Ping
Lin, Dahua
Dong, Hao
Pang, Jiangmiao
Robotics
Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representations or generative models, also referred to as world models, using large-scale visual datasets. This paper presents an end-to-end paradigm that predicts actions using inverse dynamics models conditioned on the robot's forecasted visual states, named Predictive Inverse Dynamics Models (PIDM). By closing the loop between vision and action, the end-to-end PIDM can be a better scalable action learner. In practice, we use Transformers to process both visual states and actions, naming the model Seer. It is initially pre-trained on large-scale robotic datasets, such as DROID, and can be adapted to realworld scenarios with a little fine-tuning data. Thanks to large-scale, end-to-end training and the synergy between vision and action, Seer significantly outperforms previous methods across both simulation and real-world experiments. It achieves improvements of 13% on the LIBERO-LONG benchmark, 21% on CALVIN ABC-D, and 43% in real-world tasks. Notably, Seer sets a new state-of-the-art on CALVIN ABC-D benchmark, achieving an average length of 4.28, and exhibits superior generalization for novel objects, lighting conditions, and environments under high-intensity disturbances on real-world scenarios. Code and models are publicly available at https://github.com/OpenRobotLab/Seer/.
title Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
topic Robotics
url https://arxiv.org/abs/2412.15109