Spatial RoboGrasp: Generalized Robotic Grasping Control Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yiqi, Davies, Travis, Yan, Jiahuan, Sun, Jiankai, Chen, Xiang, Hu, Luhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909624850972672
author Huang, Yiqi
Davies, Travis
Yan, Jiahuan
Sun, Jiankai
Chen, Xiang
Hu, Luhui
author_facet Huang, Yiqi
Davies, Travis
Yan, Jiahuan
Sun, Jiankai
Chen, Xiang
Hu, Luhui
contents Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their reliance on raw RGB inputs and handcrafted features often leads to overfitting and poor 3D reasoning under varied lighting, occlusion, and object conditions. In this paper, we propose a unified framework that couples robust multimodal perception with reliable grasp prediction. Our architecture fuses domain-randomized augmentation, monocular depth estimation, and a depth-aware 6-DoF Grasp Prompt into a single spatial representation for downstream action planning. Conditioned on this encoding and a high-level task prompt, our diffusion-based policy yields precise action sequences, achieving up to 40% improvement in grasp success and 45% higher task success rates under environmental variation. These results demonstrate that spatially grounded perception, paired with diffusion-based imitation learning, offers a scalable and robust solution for general-purpose robotic grasping.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20814
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
Huang, Yiqi
Davies, Travis
Yan, Jiahuan
Sun, Jiankai
Chen, Xiang
Hu, Luhui
Robotics
Computer Vision and Pattern Recognition
Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their reliance on raw RGB inputs and handcrafted features often leads to overfitting and poor 3D reasoning under varied lighting, occlusion, and object conditions. In this paper, we propose a unified framework that couples robust multimodal perception with reliable grasp prediction. Our architecture fuses domain-randomized augmentation, monocular depth estimation, and a depth-aware 6-DoF Grasp Prompt into a single spatial representation for downstream action planning. Conditioned on this encoding and a high-level task prompt, our diffusion-based policy yields precise action sequences, achieving up to 40% improvement in grasp success and 45% higher task success rates under environmental variation. These results demonstrate that spatially grounded perception, paired with diffusion-based imitation learning, offers a scalable and robust solution for general-purpose robotic grasping.
title Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20814