EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zou, Ding, Wang, Feifan, Ge, Mengyu, Fan, Siyuan, Zhang, Zongbing, Chen, Wei, Wang, Lingfeng, Hu, Zhongyou, Yan, Wenrui, Gao, Zhengwei, Wang, Hao, Jin, Weizhao, Zhang, Yu, Zhao, Hainan, Zhang, Mingliang, Xi, Xianxian, Zhang, Yaru, Li, Wenyuan, Gao, Zhengguang, Zhu, Yurui
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909865685811200
author Zou, Ding
Wang, Feifan
Ge, Mengyu
Fan, Siyuan
Zhang, Zongbing
Chen, Wei
Wang, Lingfeng
Hu, Zhongyou
Yan, Wenrui
Gao, Zhengwei
Wang, Hao
Jin, Weizhao
Zhang, Yu
Zhao, Hainan
Zhang, Mingliang
Xi, Xianxian
Zhang, Yaru
Li, Wenyuan
Gao, Zhengguang
Zhu, Yurui
author_facet Zou, Ding
Wang, Feifan
Ge, Mengyu
Fan, Siyuan
Zhang, Zongbing
Chen, Wei
Wang, Lingfeng
Hu, Zhongyou
Yan, Wenrui
Gao, Zhengwei
Wang, Hao
Jin, Weizhao
Zhang, Yu
Zhao, Hainan
Zhang, Mingliang
Xi, Xianxian
Zhang, Yaru
Li, Wenyuan
Gao, Zhengguang
Zhu, Yurui
contents The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in physical environments. However, current large language models (LLMs) and multimodal LLMs (MLLMs) for embodied tasks suffer from key limitations, including a significant gap between model design and agent requirements, an unavoidable trade-off between real-time latency and performance, and the use of unauthentic, offline evaluation metrics. To address these challenges, we propose EmbodiedBrain, a novel vision-language foundation model available in both 7B and 32B parameter sizes. Our framework features an agent-aligned data structure and employs a powerful training methodology that integrates large-scale Supervised Fine-Tuning (SFT) with Step-Augumented Group Relative Policy Optimization (Step-GRPO), which boosts long-horizon task success by integrating preceding steps as Guided Precursors. Furthermore, we incorporate a comprehensive reward system, including a Generative Reward Model (GRM) accelerated at the infrastructure level, to improve training efficiency. For enable thorough validation, we establish a three-part evaluation system encompassing General, Planning, and End-to-End Simulation Benchmarks, highlighted by the proposal and open-sourcing of a novel, challenging simulation environment. Experimental results demonstrate that EmbodiedBrain achieves superior performance across all metrics, establishing a new state-of-the-art for embodied foundation models. Towards paving the way for the next generation of generalist embodied agents, we open-source all of our data, model weight, and evaluating methods, which are available at https://zterobot.github.io/EmbodiedBrain.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
Zou, Ding
Wang, Feifan
Ge, Mengyu
Fan, Siyuan
Zhang, Zongbing
Chen, Wei
Wang, Lingfeng
Hu, Zhongyou
Yan, Wenrui
Gao, Zhengwei
Wang, Hao
Jin, Weizhao
Zhang, Yu
Zhao, Hainan
Zhang, Mingliang
Xi, Xianxian
Zhang, Yaru
Li, Wenyuan
Gao, Zhengguang
Zhu, Yurui
Computer Vision and Pattern Recognition
Robotics
The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in physical environments. However, current large language models (LLMs) and multimodal LLMs (MLLMs) for embodied tasks suffer from key limitations, including a significant gap between model design and agent requirements, an unavoidable trade-off between real-time latency and performance, and the use of unauthentic, offline evaluation metrics. To address these challenges, we propose EmbodiedBrain, a novel vision-language foundation model available in both 7B and 32B parameter sizes. Our framework features an agent-aligned data structure and employs a powerful training methodology that integrates large-scale Supervised Fine-Tuning (SFT) with Step-Augumented Group Relative Policy Optimization (Step-GRPO), which boosts long-horizon task success by integrating preceding steps as Guided Precursors. Furthermore, we incorporate a comprehensive reward system, including a Generative Reward Model (GRM) accelerated at the infrastructure level, to improve training efficiency. For enable thorough validation, we establish a three-part evaluation system encompassing General, Planning, and End-to-End Simulation Benchmarks, highlighted by the proposal and open-sourcing of a novel, challenging simulation environment. Experimental results demonstrate that EmbodiedBrain achieves superior performance across all metrics, establishing a new state-of-the-art for embodied foundation models. Towards paving the way for the next generation of generalist embodied agents, we open-source all of our data, model weight, and evaluating methods, which are available at https://zterobot.github.io/EmbodiedBrain.github.io.
title EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2510.20578