GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Weiquan, Hu, Yaoqing, Dai, Liangchen, Tang, Xu, Chen, Xingyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909052395585536
author Lin, Weiquan
Hu, Yaoqing
Dai, Liangchen
Tang, Xu
Chen, Xingyu
author_facet Lin, Weiquan
Hu, Yaoqing
Dai, Liangchen
Tang, Xu
Chen, Xingyu
contents Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities caused by self-occlusions and hand-object interactions. While introducing depth can explicitly provide spatial cues, raw sensor-captured depth maps are extensively noisy and incomplete, limiting their usefulness for fine-grained hand reconstruction. To bridge this gap, we propose GeoHand, a novel framework that unlocks high-quality geometric priors from a frozen foundational monocular geometry estimator (MoGe2). Recognizing that these priors are oriented toward general scenes, we introduce a map-level GeoAdapter to recalibrate the spatial features, specifically adapting them for detailed hand reconstruction. Furthermore, to systematically integrate these adapted priors without overwhelming intrinsic RGB appearance cues, we employ a gated cross-modal token fusion strategy. Finally, to secure precise local articulation, we design a Keypoint-Queried Iterative Refiner (KQIR) that uses projected joint locations to query geometry-aware image features for spatial correction. By combining global geometric disambiguation with local refinement in a unified pipeline, GeoHand achieves state-of-the-art performance on FreiHAND, DexYCB, and HO3Dv3, especially under severe occlusions and hand-object interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17354
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
Lin, Weiquan
Hu, Yaoqing
Dai, Liangchen
Tang, Xu
Chen, Xingyu
Computer Vision and Pattern Recognition
Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities caused by self-occlusions and hand-object interactions. While introducing depth can explicitly provide spatial cues, raw sensor-captured depth maps are extensively noisy and incomplete, limiting their usefulness for fine-grained hand reconstruction. To bridge this gap, we propose GeoHand, a novel framework that unlocks high-quality geometric priors from a frozen foundational monocular geometry estimator (MoGe2). Recognizing that these priors are oriented toward general scenes, we introduce a map-level GeoAdapter to recalibrate the spatial features, specifically adapting them for detailed hand reconstruction. Furthermore, to systematically integrate these adapted priors without overwhelming intrinsic RGB appearance cues, we employ a gated cross-modal token fusion strategy. Finally, to secure precise local articulation, we design a Keypoint-Queried Iterative Refiner (KQIR) that uses projected joint locations to query geometry-aware image features for spatial correction. By combining global geometric disambiguation with local refinement in a unified pipeline, GeoHand achieves state-of-the-art performance on FreiHAND, DexYCB, and HO3Dv3, especially under severe occlusions and hand-object interactions.
title GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17354