GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Rui, Wang, Guankun, Bai, Long, Gao, Huxin, Lai, Jiewen, Ng, Chi Kit, Wang, Jiazheng, Zhang, Fan, Ren, Hongliang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917247094620160
author Tang, Rui
Wang, Guankun
Bai, Long
Gao, Huxin
Lai, Jiewen
Ng, Chi Kit
Wang, Jiazheng
Zhang, Fan
Ren, Hongliang
author_facet Tang, Rui
Wang, Guankun
Bai, Long
Gao, Huxin
Lai, Jiewen
Ng, Chi Kit
Wang, Jiazheng
Zhang, Fan
Ren, Hongliang
contents Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate object perception and grasping, which leads to limited cross-modal fusion, redundant computation, and poor generalization in cluttered, occluded, or low-texture scenes. To address these limitations, we propose GeoLanG, an end-to-end multi-task framework built upon the CLIP architecture that unifies visual and linguistic inputs into a shared representation space for robust semantic alignment and improved generalization. To enhance target discrimination under occlusion and low-texture conditions, we explore a more effective use of depth information through the Depth-guided Geometric Module (DGGM), which converts depth into explicit geometric priors and injects them into the attention mechanism without additional computational overhead. In addition, we propose Adaptive Dense Channel Integration, which adaptively balances the contributions of multi-layer features to produce more discriminative and generalizable visual representations. Extensive experiments on the OCID-VLG dataset, as well as in both simulation and real-world hardware, demonstrate that GeoLanG enables precise and robust language-guided grasping in complex, cluttered environments, paving the way toward more reliable multimodal robotic manipulation in real-world human-centric settings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04231
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
Tang, Rui
Wang, Guankun
Bai, Long
Gao, Huxin
Lai, Jiewen
Ng, Chi Kit
Wang, Jiazheng
Zhang, Fan
Ren, Hongliang
Robotics
Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate object perception and grasping, which leads to limited cross-modal fusion, redundant computation, and poor generalization in cluttered, occluded, or low-texture scenes. To address these limitations, we propose GeoLanG, an end-to-end multi-task framework built upon the CLIP architecture that unifies visual and linguistic inputs into a shared representation space for robust semantic alignment and improved generalization. To enhance target discrimination under occlusion and low-texture conditions, we explore a more effective use of depth information through the Depth-guided Geometric Module (DGGM), which converts depth into explicit geometric priors and injects them into the attention mechanism without additional computational overhead. In addition, we propose Adaptive Dense Channel Integration, which adaptively balances the contributions of multi-layer features to produce more discriminative and generalizable visual representations. Extensive experiments on the OCID-VLG dataset, as well as in both simulation and real-world hardware, demonstrate that GeoLanG enables precise and robust language-guided grasping in complex, cluttered environments, paving the way toward more reliable multimodal robotic manipulation in real-world human-centric settings.
title GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
topic Robotics
url https://arxiv.org/abs/2602.04231