CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910251245109248 |
|---|---|
| author | Song, Kun Chen, Gaoming Ma, Shentao Jin, Ninglong Zhao, Guangbao Ding, Mingyu Xiong, Zhenhua Pan, Jia |
| author_facet | Song, Kun Chen, Gaoming Ma, Shentao Jin, Ninglong Zhao, Guangbao Ding, Mingyu Xiong, Zhenhua Pan, Jia |
| contents | One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require large-object manipulation, such as moving tables, where multiple robots must work collaboratively. Existing studies still lack a generalizable framework that can handle diverse objects, tasks, and robot team sizes. In this work, we propose CollaBot, a generalist framework for simultaneous collaborative manipulation. First, we use SEEM for scene segmentation and target-object extraction. Then, we propose a collaborative grasping framework that decomposes the task into local grasp pose generation and global coordination. Finally, we design a two-stage planning module to generate collision-free trajectories for task execution. Experimental results across different settings with varying objects, tasks, and numbers of robots indicate that our framework achieves a 72% success rate. This marks a substantial improvement over behavior cloning-based methods, validating the advantages of the proposed framework in complex multi-robot cooperative tasks. Real-world experiments further demonstrate the feasibility of our method in practical applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_03526 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation Song, Kun Chen, Gaoming Ma, Shentao Jin, Ninglong Zhao, Guangbao Ding, Mingyu Xiong, Zhenhua Pan, Jia Robotics One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require large-object manipulation, such as moving tables, where multiple robots must work collaboratively. Existing studies still lack a generalizable framework that can handle diverse objects, tasks, and robot team sizes. In this work, we propose CollaBot, a generalist framework for simultaneous collaborative manipulation. First, we use SEEM for scene segmentation and target-object extraction. Then, we propose a collaborative grasping framework that decomposes the task into local grasp pose generation and global coordination. Finally, we design a two-stage planning module to generate collision-free trajectories for task execution. Experimental results across different settings with varying objects, tasks, and numbers of robots indicate that our framework achieves a 72% success rate. This marks a substantial improvement over behavior cloning-based methods, validating the advantages of the proposed framework in complex multi-robot cooperative tasks. Real-world experiments further demonstrate the feasibility of our method in practical applications. |
| title | CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation |
| topic | Robotics |
| url | https://arxiv.org/abs/2508.03526 |