SECOND-Grasp: Semantic Contact-guided Dexterous Grasping

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shin, Han Yi, Ko, Heeju, Mun, Jaewon, Huang, Qixing, Lee, Jaehyeok, Kim, Sung June, Lee, Honglak, Jang, Sujin, Kim, Sangpil
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917490727059456
author Shin, Han Yi
Ko, Heeju
Mun, Jaewon
Huang, Qixing
Lee, Jaehyeok
Kim, Sung June
Lee, Honglak
Jang, Sujin
Kim, Sangpil
author_facet Shin, Han Yi
Ko, Heeju
Mun, Jaewon
Huang, Qixing
Lee, Jaehyeok
Kim, Sung June
Lee, Honglak
Jang, Sujin
Kim, Sangpil
contents Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guidance, yet these objectives are often treated as separate, disjoint goals. In this paper, we investigate how to integrate dexterous grasping techniques, i.e., physically stable grasps for object lifting and language-guided grasp generation, to achieve both physical stability and semantic understanding. To this end, we propose SECOND-Grasp (SEmantic CONtact-guided Dexterous Grasping), a unified framework that enables robotic hands to dynamically adjust grasping strategies based on semantic reasoning while ensuring physical feasibility. We begin by obtaining coarse contact proposals through vision-language reasoning to infer where contacts should occur based on object properties, followed by segmentation to localize these regions across views. To further ensure consistency across multiple viewpoints, we introduce Semantic-Geometric Consistency Refinement (SGCR), which refines initial contact predictions by enforcing semantic consistency across views and removing geometrically invalid regions, yielding reliable 3D contact maps. Then, we derive a feasible hand pose for each contact map via inverse kinematics, generating a supervision signal for policy learning. Our approach, trained on DexGraspNet, consistently outperforms baselines in lifting success rate on both seen and unseen categories, achieving 98.2% and 97.7%, respectively, while also improving intent-aware grasping by 12.8% and 26.2%. We further show promising results on additional datasets and robotic hands, including Shadow Hand and Allegro Hand.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13117
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SECOND-Grasp: Semantic Contact-guided Dexterous Grasping
Shin, Han Yi
Ko, Heeju
Mun, Jaewon
Huang, Qixing
Lee, Jaehyeok
Kim, Sung June
Lee, Honglak
Jang, Sujin
Kim, Sangpil
Robotics
Artificial Intelligence
Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guidance, yet these objectives are often treated as separate, disjoint goals. In this paper, we investigate how to integrate dexterous grasping techniques, i.e., physically stable grasps for object lifting and language-guided grasp generation, to achieve both physical stability and semantic understanding. To this end, we propose SECOND-Grasp (SEmantic CONtact-guided Dexterous Grasping), a unified framework that enables robotic hands to dynamically adjust grasping strategies based on semantic reasoning while ensuring physical feasibility. We begin by obtaining coarse contact proposals through vision-language reasoning to infer where contacts should occur based on object properties, followed by segmentation to localize these regions across views. To further ensure consistency across multiple viewpoints, we introduce Semantic-Geometric Consistency Refinement (SGCR), which refines initial contact predictions by enforcing semantic consistency across views and removing geometrically invalid regions, yielding reliable 3D contact maps. Then, we derive a feasible hand pose for each contact map via inverse kinematics, generating a supervision signal for policy learning. Our approach, trained on DexGraspNet, consistently outperforms baselines in lifting success rate on both seen and unseen categories, achieving 98.2% and 97.7%, respectively, while also improving intent-aware grasping by 12.8% and 26.2%. We further show promising results on additional datasets and robotic hands, including Shadow Hand and Allegro Hand.
title SECOND-Grasp: Semantic Contact-guided Dexterous Grasping
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2605.13117