D$^3$-Human: Dynamic Disentangled Digital Human from Monocular Video

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Honghu, Peng, Bo, Tao, Yunfan, Zhang, Juyong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916549823037440
author Chen, Honghu
Peng, Bo
Tao, Yunfan
Zhang, Juyong
author_facet Chen, Honghu
Peng, Bo
Tao, Yunfan
Zhang, Juyong
contents We introduce D$^3$-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed human bodies or only reconstructing clothing, making it difficult to apply directly in applications such as animation production. The challenge in reconstructing decoupled clothing and body lies in the occlusion caused by clothing over the body. To this end, the details of the visible area and the plausibility of the invisible area must be ensured during the reconstruction process. Our proposed method combines explicit and implicit representations to model the decoupled clothed human body, leveraging the robustness of explicit representations and the flexibility of implicit representations. Specifically, we reconstruct the visible region as SDF and propose a novel human manifold signed distance field (hmSDF) to segment the visible clothing and visible body, and then merge the visible and invisible body. Extensive experimental results demonstrate that, compared with existing reconstruction schemes, D$^3$-Human can achieve high-quality decoupled reconstruction of the human body wearing different clothing, and can be directly applied to clothing transfer and animation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01589
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle D$^3$-Human: Dynamic Disentangled Digital Human from Monocular Video
Chen, Honghu
Peng, Bo
Tao, Yunfan
Zhang, Juyong
Computer Vision and Pattern Recognition
Graphics
We introduce D$^3$-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed human bodies or only reconstructing clothing, making it difficult to apply directly in applications such as animation production. The challenge in reconstructing decoupled clothing and body lies in the occlusion caused by clothing over the body. To this end, the details of the visible area and the plausibility of the invisible area must be ensured during the reconstruction process. Our proposed method combines explicit and implicit representations to model the decoupled clothed human body, leveraging the robustness of explicit representations and the flexibility of implicit representations. Specifically, we reconstruct the visible region as SDF and propose a novel human manifold signed distance field (hmSDF) to segment the visible clothing and visible body, and then merge the visible and invisible body. Extensive experimental results demonstrate that, compared with existing reconstruction schemes, D$^3$-Human can achieve high-quality decoupled reconstruction of the human body wearing different clothing, and can be directly applied to clothing transfer and animation.
title D$^3$-Human: Dynamic Disentangled Digital Human from Monocular Video
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2501.01589