Obstruction reasoning for robotic grasping

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiao, Runyu, Bortolon, Matteo, Giuliari, Francesco, Fasoli, Alice, Povoli, Sergio, Mei, Guofeng, Wang, Yiming, Poiesi, Fabio
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915643374174208
author Jiao, Runyu
Bortolon, Matteo
Giuliari, Francesco
Fasoli, Alice
Povoli, Sergio
Mei, Guofeng
Wang, Yiming
Poiesi, Fabio
author_facet Jiao, Runyu
Bortolon, Matteo
Giuliari, Francesco
Fasoli, Alice
Povoli, Sergio
Mei, Guofeng
Wang, Yiming
Poiesi, Fabio
contents Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embodied reasoning models show emergent spatial understanding, they remain limited in terms of obstruction reasoning and accessibility planning. To bridge this gap, we present UNOGrasp, a learning-based vision-language model capable of performing visually-grounded obstruction reasoning to infer the sequence of actions needed to unobstruct the path and grasp the target object. We devise a novel multi-step reasoning process based on obstruction paths originated by the target object. We anchor each reasoning step with obstruction-aware visual cues to incentivize reasoning capability. UNOGrasp combines supervised and reinforcement finetuning through verifiable reasoning rewards. Moreover, we construct UNOBench, a large-scale dataset for both training and benchmarking, based on MetaGraspNetV2, with over 100k obstruction paths annotated by humans with obstruction ratios, contact points, and natural-language instructions. Extensive experiments and real-robot evaluations show that UNOGrasp significantly improves obstruction reasoning and grasp success across both synthetic and real-world environments, outperforming generalist and proprietary alternatives. Project website: https://tev-fbk.github.io/UnoGrasp/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Obstruction reasoning for robotic grasping
Jiao, Runyu
Bortolon, Matteo
Giuliari, Francesco
Fasoli, Alice
Povoli, Sergio
Mei, Guofeng
Wang, Yiming
Poiesi, Fabio
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embodied reasoning models show emergent spatial understanding, they remain limited in terms of obstruction reasoning and accessibility planning. To bridge this gap, we present UNOGrasp, a learning-based vision-language model capable of performing visually-grounded obstruction reasoning to infer the sequence of actions needed to unobstruct the path and grasp the target object. We devise a novel multi-step reasoning process based on obstruction paths originated by the target object. We anchor each reasoning step with obstruction-aware visual cues to incentivize reasoning capability. UNOGrasp combines supervised and reinforcement finetuning through verifiable reasoning rewards. Moreover, we construct UNOBench, a large-scale dataset for both training and benchmarking, based on MetaGraspNetV2, with over 100k obstruction paths annotated by humans with obstruction ratios, contact points, and natural-language instructions. Extensive experiments and real-robot evaluations show that UNOGrasp significantly improves obstruction reasoning and grasp success across both synthetic and real-world environments, outperforming generalist and proprietary alternatives. Project website: https://tev-fbk.github.io/UnoGrasp/.
title Obstruction reasoning for robotic grasping
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.23186