RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Feng, Liu, Fanfan, Zheng, Liming, Zhong, Yufeng, Huang, Yiyang, Guan, Zechao, Feng, Chengjian, Ma, Lin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918184021393408
author Yan, Feng
Liu, Fanfan
Zheng, Liming
Zhong, Yufeng
Huang, Yiyang
Guan, Zechao
Feng, Chengjian
Ma, Lin
author_facet Yan, Feng
Liu, Fanfan
Zheng, Liming
Zhong, Yufeng
Huang, Yiyang
Guan, Zechao
Feng, Chengjian
Ma, Lin
contents Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation model RoboTron-Mani and the comprehensive dataset RoboData. RoboTron-Mani, on one hand, enhances 3D perception through camera parameters and occupancy supervision. On the other hand, it further incorporates Modality-Isolation-Mask and multimodal decoder blocks based on OpenFlamingo, improving modality fusion and fine-grained perception. RoboData integrats several publicly-available datasets, achieving the first fusion of multi-view images, camera parameters, depth maps, actions, and space alignment, which facilitates comprehensive learning from diverse robotic datasets and offers one complete evaluation system. Trained on RoboData, RoboTron-Mani is the first generalist policy that surpasses expert models, enabling simultaneous evaluation of all tasks across multiple datasets, rather than being limited to specific data or task selections. Specifically, RoboTron-Mani boosts manipulation performance by increasing the average sequence length on CALVIN from 1.7 to 3.5, enabling cross-embodiment generalization, and achieving state-of-the-art results on both simulated and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07215
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
Yan, Feng
Liu, Fanfan
Zheng, Liming
Zhong, Yufeng
Huang, Yiyang
Guan, Zechao
Feng, Chengjian
Ma, Lin
Robotics
Multimedia
Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation model RoboTron-Mani and the comprehensive dataset RoboData. RoboTron-Mani, on one hand, enhances 3D perception through camera parameters and occupancy supervision. On the other hand, it further incorporates Modality-Isolation-Mask and multimodal decoder blocks based on OpenFlamingo, improving modality fusion and fine-grained perception. RoboData integrats several publicly-available datasets, achieving the first fusion of multi-view images, camera parameters, depth maps, actions, and space alignment, which facilitates comprehensive learning from diverse robotic datasets and offers one complete evaluation system. Trained on RoboData, RoboTron-Mani is the first generalist policy that surpasses expert models, enabling simultaneous evaluation of all tasks across multiple datasets, rather than being limited to specific data or task selections. Specifically, RoboTron-Mani boosts manipulation performance by increasing the average sequence length on CALVIN from 1.7 to 3.5, enabling cross-embodiment generalization, and achieving state-of-the-art results on both simulated and real-world datasets.
title RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
topic Robotics
Multimedia
url https://arxiv.org/abs/2412.07215