SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ho, Hsuan-I, Song, Jie, Hilliges, Otmar
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914735268560896
author Ho, Hsuan-I
Song, Jie
Hilliges, Otmar
author_facet Ho, Hsuan-I
Song, Jie
Hilliges, Otmar
contents A long-standing goal of 3D human reconstruction is to create lifelike and fully detailed 3D humans from single-view images. The main challenge lies in inferring unknown body shapes, appearances, and clothing details in areas not visible in the images. To address this, we propose SiTH, a novel pipeline that uniquely integrates an image-conditioned diffusion model into a 3D mesh reconstruction workflow. At the core of our method lies the decomposition of the challenging single-view reconstruction problem into generative hallucination and reconstruction subproblems. For the former, we employ a powerful generative diffusion model to hallucinate unseen back-view appearance based on the input images. For the latter, we leverage skinned body meshes as guidance to recover full-body texture meshes from the input and back-view images. SiTH requires as few as 500 3D human scans for training while maintaining its generality and robustness to diverse images. Extensive evaluations on two 3D human benchmarks, including our newly created one, highlighted our method's superior accuracy and perceptual quality in 3D textured human reconstruction. Our code and evaluation benchmark are available at https://ait.ethz.ch/sith
format Preprint
id arxiv_https___arxiv_org_abs_2311_15855
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion
Ho, Hsuan-I
Song, Jie
Hilliges, Otmar
Computer Vision and Pattern Recognition
A long-standing goal of 3D human reconstruction is to create lifelike and fully detailed 3D humans from single-view images. The main challenge lies in inferring unknown body shapes, appearances, and clothing details in areas not visible in the images. To address this, we propose SiTH, a novel pipeline that uniquely integrates an image-conditioned diffusion model into a 3D mesh reconstruction workflow. At the core of our method lies the decomposition of the challenging single-view reconstruction problem into generative hallucination and reconstruction subproblems. For the former, we employ a powerful generative diffusion model to hallucinate unseen back-view appearance based on the input images. For the latter, we leverage skinned body meshes as guidance to recover full-body texture meshes from the input and back-view images. SiTH requires as few as 500 3D human scans for training while maintaining its generality and robustness to diverse images. Extensive evaluations on two 3D human benchmarks, including our newly created one, highlighted our method's superior accuracy and perceptual quality in 3D textured human reconstruction. Our code and evaluation benchmark are available at https://ait.ethz.ch/sith
title SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.15855