StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Min, Dongchan, Song, Minyoung, Ko, Eunji, Hwang, Sung Ju
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929276751708160
author Min, Dongchan
Song, Minyoung
Ko, Eunji
Hwang, Sung Ju
author_facet Min, Dongchan
Song, Minyoung
Ko, Eunji
Hwang, Sung Ju
contents We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks. Specifically, by leveraging a pretrained image generator and an image encoder, we estimate the latent codes of the talking head video that faithfully reflects the given audio. This is made possible with several newly devised components: 1) A contrastive lip-sync discriminator for accurate lip synchronization, 2) A conditional sequential variational autoencoder that learns the latent motion space disentangled from the lip movements, such that we can independently manipulate the motions and lip movements while preserving the identity. 3) An auto-regressive prior augmented with normalizing flow to learn a complex audio-to-motion multi-modal latent space. Equipped with these components, StyleTalker can generate talking head videos not only in a motion-controllable way when another motion source video is given but also in a completely audio-driven manner by inferring realistic motions from the input audio. Through extensive experiments and user studies, we show that our model is able to synthesize talking head videos with impressive perceptual quality which are accurately lip-synced with the input audios, largely outperforming state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2208_10922
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
Min, Dongchan
Song, Minyoung
Ko, Eunji
Hwang, Sung Ju
Computer Vision and Pattern Recognition
Machine Learning
Audio and Speech Processing
Image and Video Processing
We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks. Specifically, by leveraging a pretrained image generator and an image encoder, we estimate the latent codes of the talking head video that faithfully reflects the given audio. This is made possible with several newly devised components: 1) A contrastive lip-sync discriminator for accurate lip synchronization, 2) A conditional sequential variational autoencoder that learns the latent motion space disentangled from the lip movements, such that we can independently manipulate the motions and lip movements while preserving the identity. 3) An auto-regressive prior augmented with normalizing flow to learn a complex audio-to-motion multi-modal latent space. Equipped with these components, StyleTalker can generate talking head videos not only in a motion-controllable way when another motion source video is given but also in a completely audio-driven manner by inferring realistic motions from the input audio. Through extensive experiments and user studies, we show that our model is able to synthesize talking head videos with impressive perceptual quality which are accurately lip-synced with the input audios, largely outperforming state-of-the-art baselines.
title StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
topic Computer Vision and Pattern Recognition
Machine Learning
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2208.10922