Monocular 3D Hand Pose Estimation with Implicit Camera Alignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pantazopoulos, Christos, Thermos, Spyridon, Potamianos, Gerasimos
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915394633072640
author Pantazopoulos, Christos
Thermos, Spyridon
Potamianos, Gerasimos
author_facet Pantazopoulos, Christos
Thermos, Spyridon
Potamianos, Gerasimos
contents Estimating the 3D hand articulation from a single color image is an important problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics. Apart from the absence of depth information, occlusions, articulation complexity, and the need for camera parameters knowledge pose additional challenges. In this work, we propose an optimization pipeline for estimating the 3D hand articulation from 2D keypoint input, which includes a keypoint alignment step and a fingertip loss to overcome the need to know or estimate the camera parameters. We evaluate our approach on the EgoDexter and Dexter+Object benchmarks to showcase that it performs competitively with the state-of-the-art, while also demonstrating its robustness when processing "in-the-wild" images without any prior camera knowledge. Our quantitative analysis highlights the sensitivity of the 2D keypoint estimation accuracy, despite the use of hand priors. Code is available at the project page https://cpantazop.github.io/HandRepo/
format Preprint
id arxiv_https___arxiv_org_abs_2506_11133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Monocular 3D Hand Pose Estimation with Implicit Camera Alignment
Pantazopoulos, Christos
Thermos, Spyridon
Potamianos, Gerasimos
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Image and Video Processing
Estimating the 3D hand articulation from a single color image is an important problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics. Apart from the absence of depth information, occlusions, articulation complexity, and the need for camera parameters knowledge pose additional challenges. In this work, we propose an optimization pipeline for estimating the 3D hand articulation from 2D keypoint input, which includes a keypoint alignment step and a fingertip loss to overcome the need to know or estimate the camera parameters. We evaluate our approach on the EgoDexter and Dexter+Object benchmarks to showcase that it performs competitively with the state-of-the-art, while also demonstrating its robustness when processing "in-the-wild" images without any prior camera knowledge. Our quantitative analysis highlights the sensitivity of the 2D keypoint estimation accuracy, despite the use of hand priors. Code is available at the project page https://cpantazop.github.io/HandRepo/
title Monocular 3D Hand Pose Estimation with Implicit Camera Alignment
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2506.11133