MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Song, Chentao, Zhang, He, Yuan, Haolei, Lin, Haozhe, Tao, Jianhua, Zhang, Hongwen, Yu, Tao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912989471309824
author Song, Chentao
Zhang, He
Yuan, Haolei
Lin, Haozhe
Tao, Jianhua
Zhang, Hongwen
Yu, Tao
author_facet Song, Chentao
Zhang, He
Yuan, Haolei
Lin, Haozhe
Tao, Jianhua
Zhang, Hongwen
Yu, Tao
contents We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera assumptions. Moreover, their fully coupled feature representations make it difficult to disentangle local pose from global translation, often requiring multi-stage pipelines that introduce accumulated errors. To address these challenges, we propose MetricHMR (Metric Human Mesh Recovery), which incorporates a bounding camera ray map representation to provide explicit metric cues for human reconstruction,together with a Human Mixture-of-Experts (HumanMoE) that dynamically routes image features to specialized experts, enabling the disentangled perception of local human pose and global metric position. Leveraging the recovered metric human as a geometric anchor, we further refine monocular metric depth estimation to achieve more accurate 3D alignment between humans and scenes.Comprehensive experiments demonstrate that our method achieves state-of-the-art performance on both human mesh recovery and metric human-scene reconstruction. Project Page: https://Metaverse-AI-Lab-THU.github.io/MetricHMSR.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images
Song, Chentao
Zhang, He
Yuan, Haolei
Lin, Haozhe
Tao, Jianhua
Zhang, Hongwen
Yu, Tao
Computer Vision and Pattern Recognition
We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera assumptions. Moreover, their fully coupled feature representations make it difficult to disentangle local pose from global translation, often requiring multi-stage pipelines that introduce accumulated errors. To address these challenges, we propose MetricHMR (Metric Human Mesh Recovery), which incorporates a bounding camera ray map representation to provide explicit metric cues for human reconstruction,together with a Human Mixture-of-Experts (HumanMoE) that dynamically routes image features to specialized experts, enabling the disentangled perception of local human pose and global metric position. Leveraging the recovered metric human as a geometric anchor, we further refine monocular metric depth estimation to achieve more accurate 3D alignment between humans and scenes.Comprehensive experiments demonstrate that our method achieves state-of-the-art performance on both human mesh recovery and metric human-scene reconstruction. Project Page: https://Metaverse-AI-Lab-THU.github.io/MetricHMSR.
title MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.09919