Joker: Conditional 3D Head Synthesis with Extreme Facial Expressions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Prinzler, Malte, Zakharov, Egor, Sklyarova, Vanessa, Kabadayi, Berna, Thies, Justus
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909358776909824
author Prinzler, Malte
Zakharov, Egor
Sklyarova, Vanessa
Kabadayi, Berna
Thies, Justus
author_facet Prinzler, Malte
Zakharov, Egor
Sklyarova, Vanessa
Kabadayi, Berna
Thies, Justus
contents We introduce Joker, a new method for the conditional synthesis of 3D human heads with extreme expressions. Given a single reference image of a person, we synthesize a volumetric human head with the reference identity and a new expression. We offer control over the expression via a 3D morphable model (3DMM) and textual inputs. This multi-modal conditioning signal is essential since 3DMMs alone fail to define subtle emotional changes and extreme expressions, including those involving the mouth cavity and tongue articulation. Our method is built upon a 2D diffusion-based prior that generalizes well to out-of-domain samples, such as sculptures, heavy makeup, and paintings while achieving high levels of expressiveness. To improve view consistency, we propose a new 3D distillation technique that converts predictions of our 2D prior into a neural radiance field (NeRF). Both the 2D prior and our distillation technique produce state-of-the-art results, which are confirmed by our extensive evaluations. Also, to the best of our knowledge, our method is the first to achieve view-consistent extreme tongue articulation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16395
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Joker: Conditional 3D Head Synthesis with Extreme Facial Expressions
Prinzler, Malte
Zakharov, Egor
Sklyarova, Vanessa
Kabadayi, Berna
Thies, Justus
Computer Vision and Pattern Recognition
Graphics
We introduce Joker, a new method for the conditional synthesis of 3D human heads with extreme expressions. Given a single reference image of a person, we synthesize a volumetric human head with the reference identity and a new expression. We offer control over the expression via a 3D morphable model (3DMM) and textual inputs. This multi-modal conditioning signal is essential since 3DMMs alone fail to define subtle emotional changes and extreme expressions, including those involving the mouth cavity and tongue articulation. Our method is built upon a 2D diffusion-based prior that generalizes well to out-of-domain samples, such as sculptures, heavy makeup, and paintings while achieving high levels of expressiveness. To improve view consistency, we propose a new 3D distillation technique that converts predictions of our 2D prior into a neural radiance field (NeRF). Both the 2D prior and our distillation technique produce state-of-the-art results, which are confirmed by our extensive evaluations. Also, to the best of our knowledge, our method is the first to achieve view-consistent extreme tongue articulation.
title Joker: Conditional 3D Head Synthesis with Extreme Facial Expressions
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2410.16395