RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Haifeng, Chen, Xinyi, Chen, Yilun, Li, Hao, Han, Xiaoshen, Wang, Zehan, Wang, Tai, Pang, Jiangmiao, Zhao, Zhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908343919968256
author Huang, Haifeng
Chen, Xinyi
Chen, Yilun
Li, Hao
Han, Xiaoshen
Wang, Zehan
Wang, Tai
Pang, Jiangmiao
Zhao, Zhou
author_facet Huang, Haifeng
Chen, Xinyi
Chen, Yilun
Li, Hao
Han, Xiaoshen
Wang, Zehan
Wang, Tai
Pang, Jiangmiao
Zhao, Zhou
contents Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing two key advantages: (1) effective spatial guidance that specifies target objects and placement areas while also conveying information about object shape and size, and (2) broad generalization potential driven by large-scale vision-language models pretrained on diverse grounding datasets. We introduce RoboGround, a grounding-aware robotic manipulation system that leverages grounding masks as an intermediate representation to guide policy networks in object manipulation tasks. To further explore and enhance generalization, we propose an automated pipeline for generating large-scale, simulated data with a diverse set of objects and instructions. Extensive experiments show the value of our dataset and the effectiveness of grounding masks as intermediate guidance, significantly enhancing the generalization abilities of robot policies.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21530
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
Huang, Haifeng
Chen, Xinyi
Chen, Yilun
Li, Hao
Han, Xiaoshen
Wang, Zehan
Wang, Tai
Pang, Jiangmiao
Zhao, Zhou
Robotics
Computer Vision and Pattern Recognition
Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing two key advantages: (1) effective spatial guidance that specifies target objects and placement areas while also conveying information about object shape and size, and (2) broad generalization potential driven by large-scale vision-language models pretrained on diverse grounding datasets. We introduce RoboGround, a grounding-aware robotic manipulation system that leverages grounding masks as an intermediate representation to guide policy networks in object manipulation tasks. To further explore and enhance generalization, we propose an automated pipeline for generating large-scale, simulated data with a diverse set of objects and instructions. Extensive experiments show the value of our dataset and the effectiveness of grounding masks as intermediate guidance, significantly enhancing the generalization abilities of robot policies.
title RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.21530