RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Weixin, Zhong, Weiheng, Jiang, Zhou, Fang, Dong, Zhang, Zhongyue, Lan, Zihan, Li, Haosheng, Jia, Fan, Wang, Tiancai, Fan, Haoqiang, Yoshie, Osamu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916661475409920
author Mao, Weixin
Zhong, Weiheng
Jiang, Zhou
Fang, Dong
Zhang, Zhongyue
Lan, Zihan
Li, Haosheng
Jia, Fan
Wang, Tiancai
Fan, Haoqiang
Yoshie, Osamu
author_facet Mao, Weixin
Zhong, Weiheng
Jiang, Zhou
Fang, Dong
Zhang, Zhongyue
Lan, Zihan
Li, Haosheng
Jia, Fan
Wang, Tiancai
Fan, Haoqiang
Yoshie, Osamu
contents Existing robot policies predominantly adopt the task-centric approach, requiring end-to-end task data collection. This results in limited generalization to new tasks and difficulties in pinpointing errors within long-horizon, multi-stage tasks. To address this, we propose RoboMatrix, a skill-centric hierarchical framework designed for scalable robot task planning and execution in open-world environments. RoboMatrix extracts general meta-skills from diverse complex tasks, enabling the completion of unseen tasks through skill composition. Its architecture consists of a high-level scheduling layer that utilizes large language models (LLMs) for task decomposition, an intermediate skill layer housing meta-skill models, and a low-level hardware layer for robot control. A key innovation of our work is the introduction of the first unified vision-language-action (VLA) model capable of seamlessly integrating both movement and manipulation within one model. This is achieved by combining vision and language prompts to generate discrete actions. Experimental results demonstrate that RoboMatrix achieves a 50% higher success rate than task-centric baselines when applied to unseen objects, scenes, and tasks. To advance open-world robotics research, we will open-source code, hardware designs, model weights, and datasets at https://github.com/WayneMao/RoboMatrix.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00171
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
Mao, Weixin
Zhong, Weiheng
Jiang, Zhou
Fang, Dong
Zhang, Zhongyue
Lan, Zihan
Li, Haosheng
Jia, Fan
Wang, Tiancai
Fan, Haoqiang
Yoshie, Osamu
Robotics
Computer Vision and Pattern Recognition
Existing robot policies predominantly adopt the task-centric approach, requiring end-to-end task data collection. This results in limited generalization to new tasks and difficulties in pinpointing errors within long-horizon, multi-stage tasks. To address this, we propose RoboMatrix, a skill-centric hierarchical framework designed for scalable robot task planning and execution in open-world environments. RoboMatrix extracts general meta-skills from diverse complex tasks, enabling the completion of unseen tasks through skill composition. Its architecture consists of a high-level scheduling layer that utilizes large language models (LLMs) for task decomposition, an intermediate skill layer housing meta-skill models, and a low-level hardware layer for robot control. A key innovation of our work is the introduction of the first unified vision-language-action (VLA) model capable of seamlessly integrating both movement and manipulation within one model. This is achieved by combining vision and language prompts to generate discrete actions. Experimental results demonstrate that RoboMatrix achieves a 50% higher success rate than task-centric baselines when applied to unseen objects, scenes, and tasks. To advance open-world robotics research, we will open-source code, hardware designs, model weights, and datasets at https://github.com/WayneMao/RoboMatrix.
title RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00171