A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Airale, Louis, Vaufreydaz, Dominique, Alameda-Pineda, Xavier
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917866555572224
author Airale, Louis
Vaufreydaz, Dominique
Alameda-Pineda, Xavier
author_facet Airale, Louis
Vaufreydaz, Dominique
Alameda-Pineda, Xavier
contents Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and rendering quality while the generation of natural head motion, let alone the audio-visual correlation between head motion and speech, has often been neglected.In this work, we propose a multi-scale audio-visual synchrony loss and a multi-scale autoregressive GAN to better handle short and long-term correlation between speech and the dynamics of the head and lips.In particular, we train a stack of syncer models on multimodal input pyramids and use these models as guidance in a multi-scale generator network to produce audio-aligned motion unfolding over diverse time scales.Both the pyramid of audio-visual syncers and the generative models are trained in a low-dimensional space that fully preserves dynamics cues.The experiments show significant improvements over the state-of-the-art in head motion dynamics quality and especially in multi-scale audio-visual synchrony on a collection of benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2307_03270
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
Airale, Louis
Vaufreydaz, Dominique
Alameda-Pineda, Xavier
Graphics
Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and rendering quality while the generation of natural head motion, let alone the audio-visual correlation between head motion and speech, has often been neglected.In this work, we propose a multi-scale audio-visual synchrony loss and a multi-scale autoregressive GAN to better handle short and long-term correlation between speech and the dynamics of the head and lips.In particular, we train a stack of syncer models on multimodal input pyramids and use these models as guidance in a multi-scale generator network to produce audio-aligned motion unfolding over diverse time scales.Both the pyramid of audio-visual syncers and the generative models are trained in a low-dimensional space that fully preserves dynamics cues.The experiments show significant improvements over the state-of-the-art in head motion dynamics quality and especially in multi-scale audio-visual synchrony on a collection of benchmark datasets.
title A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
topic Graphics
Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2307.03270