BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Wentao, Wang, Bowen, Zhi, Heng, Liu, Chenyu, Li, Zhe, Liu, Jian, Lin, Zengrong, Dai, Yukun, Chen, Yipeng, Yang, Wenjie, Xie, Enci, Xue, Hao, Ji, Baixu, Xu, Chen, Wang, Zhibin, Wang, Tianshi, Zhu, Lei, Shen, Heng Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911236545839104
author Tan, Wentao
Wang, Bowen
Zhi, Heng
Liu, Chenyu
Li, Zhe
Liu, Jian
Lin, Zengrong
Dai, Yukun
Chen, Yipeng
Yang, Wenjie
Xie, Enci
Xue, Hao
Ji, Baixu
Xu, Chen
Wang, Zhibin
Wang, Tianshi
Zhu, Lei
Shen, Heng Tao
author_facet Tan, Wentao
Wang, Bowen
Zhi, Heng
Liu, Chenyu
Li, Zhe
Liu, Jian
Lin, Zengrong
Dai, Yukun
Chen, Yipeng
Yang, Wenjie
Xie, Enci
Xue, Hao
Ji, Baixu
Xu, Chen
Wang, Zhibin
Wang, Tianshi
Zhu, Lei
Shen, Heng Tao
contents Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and embodiments; vision-language-action models (VLAs) produce low-level actions yet lack robust high-level embodied reasoning; and most embodied large language models (ELLMs) are constrained to digital-space with poor generalization to the physical world. Thus, unified models that operate seamlessly across digital and physical spaces while generalizing across embodiments and tasks remain absent. We introduce the \textbf{Boundless Large Model (BLM$_1$)}, a multimodal spatial foundation model that preserves instruction following and reasoning, incorporates embodied knowledge, and supports robust cross-embodiment control. BLM$_1$ integrates three key capabilities -- \textit{cross-space transfer, cross-task learning, and cross-embodiment generalization} -- via a two-stage training paradigm. Stage I injects embodied knowledge into the MLLM through curated digital corpora while maintaining language competence. Stage II trains a policy module through an intent-bridging interface that extracts high-level semantics from the MLLM to guide control, without fine-tuning the MLLM backbone. This process is supported by a self-collected cross-embodiment demonstration suite spanning four robot embodiments and six progressively challenging tasks. Evaluations across digital and physical benchmarks show that a single BLM$_1$ instance outperforms four model families -- MLLMs, ELLMs, VLAs, and GMLMs -- achieving $\sim\!\textbf{6%}$ gains in digital tasks and $\sim\!\textbf{3%}$ in physical tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24161
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
Tan, Wentao
Wang, Bowen
Zhi, Heng
Liu, Chenyu
Li, Zhe
Liu, Jian
Lin, Zengrong
Dai, Yukun
Chen, Yipeng
Yang, Wenjie
Xie, Enci
Xue, Hao
Ji, Baixu
Xu, Chen
Wang, Zhibin
Wang, Tianshi
Zhu, Lei
Shen, Heng Tao
Artificial Intelligence
Multimedia
Robotics
Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and embodiments; vision-language-action models (VLAs) produce low-level actions yet lack robust high-level embodied reasoning; and most embodied large language models (ELLMs) are constrained to digital-space with poor generalization to the physical world. Thus, unified models that operate seamlessly across digital and physical spaces while generalizing across embodiments and tasks remain absent. We introduce the \textbf{Boundless Large Model (BLM$_1$)}, a multimodal spatial foundation model that preserves instruction following and reasoning, incorporates embodied knowledge, and supports robust cross-embodiment control. BLM$_1$ integrates three key capabilities -- \textit{cross-space transfer, cross-task learning, and cross-embodiment generalization} -- via a two-stage training paradigm. Stage I injects embodied knowledge into the MLLM through curated digital corpora while maintaining language competence. Stage II trains a policy module through an intent-bridging interface that extracts high-level semantics from the MLLM to guide control, without fine-tuning the MLLM backbone. This process is supported by a self-collected cross-embodiment demonstration suite spanning four robot embodiments and six progressively challenging tasks. Evaluations across digital and physical benchmarks show that a single BLM$_1$ instance outperforms four model families -- MLLMs, ELLMs, VLAs, and GMLMs -- achieving $\sim\!\textbf{6%}$ gains in digital tasks and $\sim\!\textbf{3%}$ in physical tasks.
title BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
topic Artificial Intelligence
Multimedia
Robotics
url https://arxiv.org/abs/2510.24161