EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Zibin, Ni, Fei, Yuan, Yifu, Li, Yinchuan, Hao, Jianye
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913839099936768
author Dong, Zibin
Ni, Fei
Yuan, Yifu
Li, Yinchuan
Hao, Jianye
author_facet Dong, Zibin
Ni, Fei
Yuan, Yifu
Li, Yinchuan
Hao, Jianye
contents We present EmbodiedMAE, a unified 3D multi-modal representation for robot manipulation. Current approaches suffer from significant domain gaps between training datasets and robot manipulation tasks, while also lacking model architectures that can effectively incorporate 3D information. To overcome these limitations, we enhance the DROID dataset with high-quality depth maps and point clouds, constructing DROID-3D as a valuable supplement for 3D embodied vision research. Then we develop EmbodiedMAE, a multi-modal masked autoencoder that simultaneously learns representations across RGB, depth, and point cloud modalities through stochastic masking and cross-modal fusion. Trained on DROID-3D, EmbodiedMAE consistently outperforms state-of-the-art vision foundation models (VFMs) in both training efficiency and final performance across 70 simulation tasks and 20 real-world robot manipulation tasks on two robot platforms. The model exhibits strong scaling behavior with size and promotes effective policy learning from 3D inputs. Experimental results establish EmbodiedMAE as a reliable unified 3D multi-modal VFM for embodied AI systems, particularly in precise tabletop manipulation settings where spatial perception is critical.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
Dong, Zibin
Ni, Fei
Yuan, Yifu
Li, Yinchuan
Hao, Jianye
Robotics
Artificial Intelligence
We present EmbodiedMAE, a unified 3D multi-modal representation for robot manipulation. Current approaches suffer from significant domain gaps between training datasets and robot manipulation tasks, while also lacking model architectures that can effectively incorporate 3D information. To overcome these limitations, we enhance the DROID dataset with high-quality depth maps and point clouds, constructing DROID-3D as a valuable supplement for 3D embodied vision research. Then we develop EmbodiedMAE, a multi-modal masked autoencoder that simultaneously learns representations across RGB, depth, and point cloud modalities through stochastic masking and cross-modal fusion. Trained on DROID-3D, EmbodiedMAE consistently outperforms state-of-the-art vision foundation models (VFMs) in both training efficiency and final performance across 70 simulation tasks and 20 real-world robot manipulation tasks on two robot platforms. The model exhibits strong scaling behavior with size and promotes effective policy learning from 3D inputs. Experimental results establish EmbodiedMAE as a reliable unified 3D multi-modal VFM for embodied AI systems, particularly in precise tabletop manipulation settings where spatial perception is critical.
title EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2505.10105