HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qi, Haozhe, Zhao, Chen, Salzmann, Mathieu, Mathis, Alexander
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917598953734144
author Qi, Haozhe
Zhao, Chen
Salzmann, Mathieu
Mathis, Alexander
author_facet Qi, Haozhe
Zhao, Chen
Salzmann, Mathieu
Mathis, Alexander
contents Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus, existing methods often rely on intermediate 3D shape representations to increase performance. These representations are typically explicit, such as 3D point clouds or meshes, and thus provide information in the direct surroundings of the intermediate hand pose estimate. To address this, we introduce HOISDF, a Signed Distance Field (SDF) guided hand-object pose estimation network, which jointly exploits hand and object SDFs to provide a global, implicit representation over the complete reconstruction volume. Specifically, the role of the SDFs is threefold: equip the visual encoder with implicit shape information, help to encode hand-object interactions, and guide the hand and object pose regression via SDF-based sampling and by augmenting the feature representations. We show that HOISDF achieves state-of-the-art results on hand-object pose estimation benchmarks (DexYCB and HO3Dv2). Code is available at https://github.com/amathislab/HOISDF
format Preprint
id arxiv_https___arxiv_org_abs_2402_17062
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields
Qi, Haozhe
Zhao, Chen
Salzmann, Mathieu
Mathis, Alexander
Computer Vision and Pattern Recognition
Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus, existing methods often rely on intermediate 3D shape representations to increase performance. These representations are typically explicit, such as 3D point clouds or meshes, and thus provide information in the direct surroundings of the intermediate hand pose estimate. To address this, we introduce HOISDF, a Signed Distance Field (SDF) guided hand-object pose estimation network, which jointly exploits hand and object SDFs to provide a global, implicit representation over the complete reconstruction volume. Specifically, the role of the SDFs is threefold: equip the visual encoder with implicit shape information, help to encode hand-object interactions, and guide the hand and object pose regression via SDF-based sampling and by augmenting the feature representations. We show that HOISDF achieves state-of-the-art results on hand-object pose estimation benchmarks (DexYCB and HO3Dv2). Code is available at https://github.com/amathislab/HOISDF
title HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.17062