Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Gen, Yang, Ganlin, Gong, Ziyang, Chen, Guanzhou, Duan, Haonan, Cui, Erfei, Tong, Ronglei, Hou, Zhi, Zhang, Tianyi, Chen, Zhe, Ye, Shenglong, Lu, Lewei, Wang, Jingbo, Wang, Wenhai, Dai, Jifeng, Qiao, Yu, Ji, Rongrong, Zhu, Xizhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916770312355840
author Luo, Gen
Yang, Ganlin
Gong, Ziyang
Chen, Guanzhou
Duan, Haonan
Cui, Erfei
Tong, Ronglei
Hou, Zhi
Zhang, Tianyi
Chen, Zhe
Ye, Shenglong
Lu, Lewei
Wang, Jingbo
Wang, Wenhai
Dai, Jifeng
Qiao, Yu
Ji, Rongrong
Zhu, Xizhou
author_facet Luo, Gen
Yang, Ganlin
Gong, Ziyang
Chen, Guanzhou
Duan, Haonan
Cui, Erfei
Tong, Ronglei
Hou, Zhi
Zhang, Tianyi
Chen, Zhe
Ye, Shenglong
Lu, Lewei
Wang, Jingbo
Wang, Wenhai
Dai, Jifeng
Qiao, Yu
Ji, Rongrong
Zhu, Xizhou
contents The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding abilities, but also integrate visual-spatial reasoning and physical interaction capabilities. Nevertheless,existing methods struggle to unify these capabilities due to their fundamental differences.In this paper, we present the Visual Embodied Brain (VeBrain), a unified framework for perception, reasoning, and control in real world. VeBrain reformulates robotic control into common text-based MLLM tasks in the 2D visual space, thus unifying the objectives and mapping spaces of different tasks. Then, a novel robotic adapter is proposed to convert textual control signals from MLLMs to motion policies of real robots. From the data perspective, we further introduce VeBrain-600k, a high-quality instruction dataset encompassing various capabilities of VeBrain. In VeBrain-600k, we take hundreds of hours to collect, curate and annotate the data, and adopt multimodal chain-of-thought(CoT) to mix the different capabilities into a single conversation. Extensive experiments on 13 multimodal benchmarks and 5 spatial intelligence benchmarks demonstrate the superior performance of VeBrain to existing MLLMs like Qwen2.5-VL. When deployed to legged robots and robotic arms, VeBrain shows strong adaptability, flexibility, and compositional capabilities compared to existing methods. For example, compared to Qwen2.5-VL, VeBrain not only achieves substantial gains on MMVet by +5.6%, but also excels in legged robot tasks with +50% average gains.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00123
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
Luo, Gen
Yang, Ganlin
Gong, Ziyang
Chen, Guanzhou
Duan, Haonan
Cui, Erfei
Tong, Ronglei
Hou, Zhi
Zhang, Tianyi
Chen, Zhe
Ye, Shenglong
Lu, Lewei
Wang, Jingbo
Wang, Wenhai
Dai, Jifeng
Qiao, Yu
Ji, Rongrong
Zhu, Xizhou
Computer Vision and Pattern Recognition
Robotics
The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding abilities, but also integrate visual-spatial reasoning and physical interaction capabilities. Nevertheless,existing methods struggle to unify these capabilities due to their fundamental differences.In this paper, we present the Visual Embodied Brain (VeBrain), a unified framework for perception, reasoning, and control in real world. VeBrain reformulates robotic control into common text-based MLLM tasks in the 2D visual space, thus unifying the objectives and mapping spaces of different tasks. Then, a novel robotic adapter is proposed to convert textual control signals from MLLMs to motion policies of real robots. From the data perspective, we further introduce VeBrain-600k, a high-quality instruction dataset encompassing various capabilities of VeBrain. In VeBrain-600k, we take hundreds of hours to collect, curate and annotate the data, and adopt multimodal chain-of-thought(CoT) to mix the different capabilities into a single conversation. Extensive experiments on 13 multimodal benchmarks and 5 spatial intelligence benchmarks demonstrate the superior performance of VeBrain to existing MLLMs like Qwen2.5-VL. When deployed to legged robots and robotic arms, VeBrain shows strong adaptability, flexibility, and compositional capabilities compared to existing methods. For example, compared to Qwen2.5-VL, VeBrain not only achieves substantial gains on MMVet by +5.6%, but also excels in legged robot tasks with +50% average gains.
title Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2506.00123