JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chakkera, Sai Tanmay Reddy, Chatziagapi, Aggelina, Samaras, Dimitris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929504244465664
author Chakkera, Sai Tanmay Reddy
Chatziagapi, Aggelina
Samaras, Dimitris
author_facet Chakkera, Sai Tanmay Reddy
Chatziagapi, Aggelina
Samaras, Dimitris
contents We introduce a novel method for joint expression and audio-guided talking face generation. Recent approaches either struggle to preserve the speaker identity or fail to produce faithful facial expressions. To address these challenges, we propose a NeRF-based network. Since we train our network on monocular videos without any ground truth, it is essential to learn disentangled representations for audio and expression. We first learn audio features in a self-supervised manner, given utterances from multiple subjects. By incorporating a contrastive learning technique, we ensure that the learned audio features are aligned to the lip motion and disentangled from the muscle motion of the rest of the face. We then devise a transformer-based architecture that learns expression features, capturing long-range facial expressions and disentangling them from the speech-specific mouth movements. Through quantitative and qualitative evaluation, we demonstrate that our method can synthesize high-fidelity talking face videos, achieving state-of-the-art facial expression transfer along with lip synchronization to unseen audio.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12156
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
Chakkera, Sai Tanmay Reddy
Chatziagapi, Aggelina
Samaras, Dimitris
Computer Vision and Pattern Recognition
We introduce a novel method for joint expression and audio-guided talking face generation. Recent approaches either struggle to preserve the speaker identity or fail to produce faithful facial expressions. To address these challenges, we propose a NeRF-based network. Since we train our network on monocular videos without any ground truth, it is essential to learn disentangled representations for audio and expression. We first learn audio features in a self-supervised manner, given utterances from multiple subjects. By incorporating a contrastive learning technique, we ensure that the learned audio features are aligned to the lip motion and disentangled from the muscle motion of the rest of the face. We then devise a transformer-based architecture that learns expression features, capturing long-range facial expressions and disentangling them from the speech-specific mouth movements. Through quantitative and qualitative evaluation, we demonstrate that our method can synthesize high-fidelity talking face videos, achieving state-of-the-art facial expression transfer along with lip synchronization to unseen audio.
title JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.12156