DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Junha, Park, Eunha, Cho, Minsu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910164856078336
author Lee, Junha
Park, Eunha
Cho, Minsu
author_facet Lee, Junha
Park, Eunha
Cho, Minsu
contents Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions. While vision-language models have been applied to this problem, existing approaches directly map observations to grasp parameters without intermediate reasoning about physical interactions. We present DextER, Dexterous Grasp Generation with Embodied Reasoning, which introduces contact-based embodied reasoning for multi-finger manipulation. Our key insight is that predicting which hand links contact where on the object surface provides an embodiment-aware intermediate representation, bridging task semantics with physical constraints. DextER autoregressively generates embodied contact tokens specifying which finger links contact where on the object surface, followed by grasp tokens encoding the hand configuration. On DexGYS, DextER achieves 67.14% success rate, outperforming state-of-the-art by 3.83 p.p. with 96.4% improvement in intention alignment. We also demonstrate steerable generation through partial contact specification, providing fine-grained control over grasp synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16046
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
Lee, Junha
Park, Eunha
Cho, Minsu
Robotics
Computer Vision and Pattern Recognition
Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions. While vision-language models have been applied to this problem, existing approaches directly map observations to grasp parameters without intermediate reasoning about physical interactions. We present DextER, Dexterous Grasp Generation with Embodied Reasoning, which introduces contact-based embodied reasoning for multi-finger manipulation. Our key insight is that predicting which hand links contact where on the object surface provides an embodiment-aware intermediate representation, bridging task semantics with physical constraints. DextER autoregressively generates embodied contact tokens specifying which finger links contact where on the object surface, followed by grasp tokens encoding the hand configuration. On DexGYS, DextER achieves 67.14% success rate, outperforming state-of-the-art by 3.83 p.p. with 96.4% improvement in intention alignment. We also demonstrate steerable generation through partial contact specification, providing fine-grained control over grasp synthesis.
title DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.16046