M3DHMR: Monocular 3D Hand Mesh Recovery

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Yihong, Wu, Xianjia, Wang, Xilai, Hu, Jianqiao, Lei, Songju, Li, Xiandong, Kang, Wenxiong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918232842043392
author Lin, Yihong
Wu, Xianjia
Wang, Xilai
Hu, Jianqiao
Lei, Songju
Li, Xiandong
Kang, Wenxiong
author_facet Lin, Yihong
Wu, Xianjia
Wang, Xilai
Hu, Jianqiao
Lei, Songju
Li, Xiandong
Kang, Wenxiong
contents Monocular 3D hand mesh recovery is challenging due to high degrees of freedom of hands, 2D-to-3D ambiguity and self-occlusion. Most existing methods are either inefficient or less straightforward for predicting the position of 3D mesh vertices. Thus, we propose a new pipeline called Monocular 3D Hand Mesh Recovery (M3DHMR) to directly estimate the positions of hand mesh vertices. M3DHMR provides 2D cues for 3D tasks from a single image and uses a new spiral decoder consist of several Dynamic Spiral Convolution (DSC) Layers and a Region of Interest (ROI) Layer. On the one hand, DSC Layers adaptively adjust the weights based on the vertex positions and extract the vertex features in both spatial and channel dimensions. On the other hand, ROI Layer utilizes the physical information and refines mesh vertices in each predefined hand region separately. Extensive experiments on popular dataset FreiHAND demonstrate that M3DHMR significantly outperforms state-of-the-art real-time methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20058
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M3DHMR: Monocular 3D Hand Mesh Recovery
Lin, Yihong
Wu, Xianjia
Wang, Xilai
Hu, Jianqiao
Lei, Songju
Li, Xiandong
Kang, Wenxiong
Computer Vision and Pattern Recognition
Monocular 3D hand mesh recovery is challenging due to high degrees of freedom of hands, 2D-to-3D ambiguity and self-occlusion. Most existing methods are either inefficient or less straightforward for predicting the position of 3D mesh vertices. Thus, we propose a new pipeline called Monocular 3D Hand Mesh Recovery (M3DHMR) to directly estimate the positions of hand mesh vertices. M3DHMR provides 2D cues for 3D tasks from a single image and uses a new spiral decoder consist of several Dynamic Spiral Convolution (DSC) Layers and a Region of Interest (ROI) Layer. On the one hand, DSC Layers adaptively adjust the weights based on the vertex positions and extract the vertex features in both spatial and channel dimensions. On the other hand, ROI Layer utilizes the physical information and refines mesh vertices in each predefined hand region separately. Extensive experiments on popular dataset FreiHAND demonstrate that M3DHMR significantly outperforms state-of-the-art real-time methods.
title M3DHMR: Monocular 3D Hand Mesh Recovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20058