DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shen, Wenhao, Zhou, Ming, Zhang, Hengyuan, Bian, Siyuan, Xu, Youjiang, Lin, Xi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917517921878016
author Shen, Wenhao
Zhou, Ming
Zhang, Hengyuan
Bian, Siyuan
Xu, Youjiang
Lin, Xi
author_facet Shen, Wenhao
Zhou, Ming
Zhang, Hengyuan
Bian, Siyuan
Xu, Youjiang
Lin, Xi
contents Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expressive whole-body motion are required. Existing video HMR methods produce coherent body motion but often overlook detailed hand articulation, while image-based whole-body methods recover SMPL-X meshes independently per frame, often leading to jittery and inaccurate hand motion. We present a temporally coherent whole-body HMR framework for challenging in-the-wild monocular videos. Our model unifies body context and part-specific hand observations through residual body-hand fusion, enabling stable body motion and detailed hand recovery within a single temporal architecture. We further introduce close-up-aware augmentation to improve robustness under upper-body framing. Experiments on whole-body and body-only benchmarks demonstrate improved hand reconstruction and competitive body accuracy. Our method also produces temporally stable and 2D-consistent SMPL-X motion in challenging real-world videos.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos
Shen, Wenhao
Zhou, Ming
Zhang, Hengyuan
Bian, Siyuan
Xu, Youjiang
Lin, Xi
Computer Vision and Pattern Recognition
Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expressive whole-body motion are required. Existing video HMR methods produce coherent body motion but often overlook detailed hand articulation, while image-based whole-body methods recover SMPL-X meshes independently per frame, often leading to jittery and inaccurate hand motion. We present a temporally coherent whole-body HMR framework for challenging in-the-wild monocular videos. Our model unifies body context and part-specific hand observations through residual body-hand fusion, enabling stable body motion and detailed hand recovery within a single temporal architecture. We further introduce close-up-aware augmentation to improve robustness under upper-body framing. Experiments on whole-body and body-only benchmarks demonstrate improved hand reconstruction and competitive body accuracy. Our method also produces temporally stable and 2D-consistent SMPL-X motion in challenging real-world videos.
title DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18102