FitDiff: Robust monocular 3D facial shape and reflectance estimation using Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Galanakis, Stathis, Lattas, Alexandros, Moschoglou, Stylianos, Zafeiriou, Stefanos
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916636997451776
author Galanakis, Stathis
Lattas, Alexandros
Moschoglou, Stylianos
Zafeiriou, Stefanos
author_facet Galanakis, Stathis
Lattas, Alexandros
Moschoglou, Stylianos
Zafeiriou, Stefanos
contents The remarkable progress in 3D face reconstruction has resulted in high-detail and photorealistic facial representations. Recently, Diffusion Models have revolutionized the capabilities of generative methods by surpassing the performance of GANs. In this work, we present FitDiff, a diffusion-based 3D facial avatar generative model. Leveraging diffusion principles, our model accurately generates relightable facial avatars, utilizing an identity embedding extracted from an "in-the-wild" 2D facial image. The introduced multi-modal diffusion model is the first to concurrently output facial reflectance maps (diffuse and specular albedo and normals) and shapes, showcasing great generalization capabilities. It is solely trained on an annotated subset of a public facial dataset, paired with 3D reconstructions. We revisit the typical 3D facial fitting approach by guiding a reverse diffusion process using perceptual and face recognition losses. Being the first 3D LDM conditioned on face recognition embeddings, FitDiff reconstructs relightable human avatars, that can be used as-is in common rendering engines, starting only from an unconstrained facial image, and achieving state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2312_04465
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FitDiff: Robust monocular 3D facial shape and reflectance estimation using Diffusion Models
Galanakis, Stathis
Lattas, Alexandros
Moschoglou, Stylianos
Zafeiriou, Stefanos
Computer Vision and Pattern Recognition
The remarkable progress in 3D face reconstruction has resulted in high-detail and photorealistic facial representations. Recently, Diffusion Models have revolutionized the capabilities of generative methods by surpassing the performance of GANs. In this work, we present FitDiff, a diffusion-based 3D facial avatar generative model. Leveraging diffusion principles, our model accurately generates relightable facial avatars, utilizing an identity embedding extracted from an "in-the-wild" 2D facial image. The introduced multi-modal diffusion model is the first to concurrently output facial reflectance maps (diffuse and specular albedo and normals) and shapes, showcasing great generalization capabilities. It is solely trained on an annotated subset of a public facial dataset, paired with 3D reconstructions. We revisit the typical 3D facial fitting approach by guiding a reverse diffusion process using perceptual and face recognition losses. Being the first 3D LDM conditioned on face recognition embeddings, FitDiff reconstructs relightable human avatars, that can be used as-is in common rendering engines, starting only from an unconstrained facial image, and achieving state-of-the-art performance.
title FitDiff: Robust monocular 3D facial shape and reflectance estimation using Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.04465