CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Kun, Chen, Gaoming, Ma, Shentao, Jin, Ninglong, Zhao, Guangbao, Ding, Mingyu, Xiong, Zhenhua, Pan, Jia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910251245109248
author Song, Kun
Chen, Gaoming
Ma, Shentao
Jin, Ninglong
Zhao, Guangbao
Ding, Mingyu
Xiong, Zhenhua
Pan, Jia
author_facet Song, Kun
Chen, Gaoming
Ma, Shentao
Jin, Ninglong
Zhao, Guangbao
Ding, Mingyu
Xiong, Zhenhua
Pan, Jia
contents One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require large-object manipulation, such as moving tables, where multiple robots must work collaboratively. Existing studies still lack a generalizable framework that can handle diverse objects, tasks, and robot team sizes. In this work, we propose CollaBot, a generalist framework for simultaneous collaborative manipulation. First, we use SEEM for scene segmentation and target-object extraction. Then, we propose a collaborative grasping framework that decomposes the task into local grasp pose generation and global coordination. Finally, we design a two-stage planning module to generate collision-free trajectories for task execution. Experimental results across different settings with varying objects, tasks, and numbers of robots indicate that our framework achieves a 72% success rate. This marks a substantial improvement over behavior cloning-based methods, validating the advantages of the proposed framework in complex multi-robot cooperative tasks. Real-world experiments further demonstrate the feasibility of our method in practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03526
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
Song, Kun
Chen, Gaoming
Ma, Shentao
Jin, Ninglong
Zhao, Guangbao
Ding, Mingyu
Xiong, Zhenhua
Pan, Jia
Robotics
One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require large-object manipulation, such as moving tables, where multiple robots must work collaboratively. Existing studies still lack a generalizable framework that can handle diverse objects, tasks, and robot team sizes. In this work, we propose CollaBot, a generalist framework for simultaneous collaborative manipulation. First, we use SEEM for scene segmentation and target-object extraction. Then, we propose a collaborative grasping framework that decomposes the task into local grasp pose generation and global coordination. Finally, we design a two-stage planning module to generate collision-free trajectories for task execution. Experimental results across different settings with varying objects, tasks, and numbers of robots indicate that our framework achieves a 72% success rate. This marks a substantial improvement over behavior cloning-based methods, validating the advantages of the proposed framework in complex multi-robot cooperative tasks. Real-world experiments further demonstrate the feasibility of our method in practical applications.
title CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
topic Robotics
url https://arxiv.org/abs/2508.03526