DiffusionAvatars: Deferred Diffusion for High-fidelity 3D Head Avatars

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kirschstein, Tobias, Giebenhain, Simon, Nießner, Matthias
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916208664641536
author Kirschstein, Tobias
Giebenhain, Simon
Nießner, Matthias
author_facet Kirschstein, Tobias
Giebenhain, Simon
Nießner, Matthias
contents DiffusionAvatars synthesizes a high-fidelity 3D head avatar of a person, offering intuitive control over both pose and expression. We propose a diffusion-based neural renderer that leverages generic 2D priors to produce compelling images of faces. For coarse guidance of the expression and head pose, we render a neural parametric head model (NPHM) from the target viewpoint, which acts as a proxy geometry of the person. Additionally, to enhance the modeling of intricate facial expressions, we condition DiffusionAvatars directly on the expression codes obtained from NPHM via cross-attention. Finally, to synthesize consistent surface details across different viewpoints and expressions, we rig learnable spatial features to the head's surface via TriPlane lookup in NPHM's canonical space. We train DiffusionAvatars on RGB videos and corresponding fitted NPHM meshes of a person and test the obtained avatars in both self-reenactment and animation scenarios. Our experiments demonstrate that DiffusionAvatars generates temporally consistent and visually appealing videos for novel poses and expressions of a person, outperforming existing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18635
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DiffusionAvatars: Deferred Diffusion for High-fidelity 3D Head Avatars
Kirschstein, Tobias
Giebenhain, Simon
Nießner, Matthias
Computer Vision and Pattern Recognition
DiffusionAvatars synthesizes a high-fidelity 3D head avatar of a person, offering intuitive control over both pose and expression. We propose a diffusion-based neural renderer that leverages generic 2D priors to produce compelling images of faces. For coarse guidance of the expression and head pose, we render a neural parametric head model (NPHM) from the target viewpoint, which acts as a proxy geometry of the person. Additionally, to enhance the modeling of intricate facial expressions, we condition DiffusionAvatars directly on the expression codes obtained from NPHM via cross-attention. Finally, to synthesize consistent surface details across different viewpoints and expressions, we rig learnable spatial features to the head's surface via TriPlane lookup in NPHM's canonical space. We train DiffusionAvatars on RGB videos and corresponding fitted NPHM meshes of a person and test the obtained avatars in both self-reenactment and animation scenarios. Our experiments demonstrate that DiffusionAvatars generates temporally consistent and visually appealing videos for novel poses and expressions of a person, outperforming existing approaches.
title DiffusionAvatars: Deferred Diffusion for High-fidelity 3D Head Avatars
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.18635