Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Yu, Zhang, Shunsi, Shu, Jian, Zhao, Hanfeng, Pang, Guoliang, Zhang, Chi, Wang, Hao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915047811317760
author Feng, Yu
Zhang, Shunsi
Shu, Jian
Zhao, Hanfeng
Pang, Guoliang
Zhang, Chi
Wang, Hao
author_facet Feng, Yu
Zhang, Shunsi
Shu, Jian
Zhao, Hanfeng
Pang, Guoliang
Zhang, Chi
Wang, Hao
contents Generating multi-view human images from a single view is a complex and significant challenge. Although recent advancements in multi-view object generation have shown impressive results with diffusion models, novel view synthesis for humans remains constrained by the limited availability of 3D human datasets. Consequently, many existing models struggle to produce realistic human body shapes or capture fine-grained facial details accurately. To address these issues, we propose an innovative framework that leverages transferred body and facial representations for multi-view human synthesis. Specifically, we use a single-view model pretrained on a large-scale human dataset to develop a multi-view body representation, aiming to extend the 2D knowledge of the single-view model to a multi-view diffusion model. Additionally, to enhance the model's detail restoration capability, we integrate transferred multimodal facial features into our trained human diffusion model. Experimental evaluations on benchmark datasets demonstrate that our approach outperforms the current state-of-the-art methods, achieving superior performance in multi-view human synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03011
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations
Feng, Yu
Zhang, Shunsi
Shu, Jian
Zhao, Hanfeng
Pang, Guoliang
Zhang, Chi
Wang, Hao
Computer Vision and Pattern Recognition
Artificial Intelligence
Generating multi-view human images from a single view is a complex and significant challenge. Although recent advancements in multi-view object generation have shown impressive results with diffusion models, novel view synthesis for humans remains constrained by the limited availability of 3D human datasets. Consequently, many existing models struggle to produce realistic human body shapes or capture fine-grained facial details accurately. To address these issues, we propose an innovative framework that leverages transferred body and facial representations for multi-view human synthesis. Specifically, we use a single-view model pretrained on a large-scale human dataset to develop a multi-view body representation, aiming to extend the 2D knowledge of the single-view model to a multi-view diffusion model. Additionally, to enhance the model's detail restoration capability, we integrate transferred multimodal facial features into our trained human diffusion model. Experimental evaluations on benchmark datasets demonstrate that our approach outperforms the current state-of-the-art methods, achieving superior performance in multi-view human synthesis.
title Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.03011