MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Yushuo, Zheng, Zerong, Li, Zhe, Xu, Chao, Liu, Yebin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911952800841728
author Chen, Yushuo
Zheng, Zerong
Li, Zhe
Xu, Chao
Liu, Yebin
author_facet Chen, Yushuo
Zheng, Zerong
Li, Zhe
Xu, Chao
Liu, Yebin
contents We present a novel pipeline for learning high-quality triangular human avatars from multi-view videos. Recent methods for avatar learning are typically based on neural radiance fields (NeRF), which is not compatible with traditional graphics pipeline and poses great challenges for operations like editing or synthesizing under different environments. To overcome these limitations, our method represents the avatar with an explicit triangular mesh extracted from an implicit SDF field, complemented by an implicit material field conditioned on given poses. Leveraging this triangular avatar representation, we incorporate physics-based rendering to accurately decompose geometry and texture. To enhance both the geometric and appearance details, we further employ a 2D UNet as the network backbone and introduce pseudo normal ground-truth as additional supervision. Experiments show that our method can learn triangular avatars with high-quality geometry reconstruction and plausible material decomposition, inherently supporting editing, manipulation or relighting operations.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08414
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos
Chen, Yushuo
Zheng, Zerong
Li, Zhe
Xu, Chao
Liu, Yebin
Computer Vision and Pattern Recognition
Graphics
We present a novel pipeline for learning high-quality triangular human avatars from multi-view videos. Recent methods for avatar learning are typically based on neural radiance fields (NeRF), which is not compatible with traditional graphics pipeline and poses great challenges for operations like editing or synthesizing under different environments. To overcome these limitations, our method represents the avatar with an explicit triangular mesh extracted from an implicit SDF field, complemented by an implicit material field conditioned on given poses. Leveraging this triangular avatar representation, we incorporate physics-based rendering to accurately decompose geometry and texture. To enhance both the geometric and appearance details, we further employ a 2D UNet as the network backbone and introduce pseudo normal ground-truth as additional supervision. Experiments show that our method can learn triangular avatars with high-quality geometry reconstruction and plausible material decomposition, inherently supporting editing, manipulation or relighting operations.
title MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2407.08414