Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Se Jin, Kim, Minsu, Choi, Jeongsoo, Ro, Yong Man
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916184801148928
author Park, Se Jin
Kim, Minsu
Choi, Jeongsoo
Ro, Yong Man
author_facet Park, Se Jin
Kim, Minsu
Choi, Jeongsoo
Ro, Yong Man
contents Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or following phones, the articulation of a phone varies upon the phonetic context. Therefore, modeling lip motion with the phonetic context can generate more spatio-temporally aligned lip movement. In this respect, we investigate the phonetic context in generating lip motion for talking face generation. We propose Context-Aware Lip-Sync framework (CALS), which explicitly leverages phonetic context to generate lip movement of the target face. CALS is comprised of an Audio-to-Lip module and a Lip-to-Face module. The former is pretrained based on masked learning to map each phone to a contextualized lip motion unit. The contextualized lip motion unit then guides the latter in synthesizing a target identity with context-aware lip motion. From extensive experiments, we verify that simply exploiting the phonetic context in the proposed CALS framework effectively enhances spatio-temporal alignment. We also demonstrate the extent to which the phonetic context assists in lip synchronization and find the effective window size for lip generation to be approximately 1.2 seconds.
format Preprint
id arxiv_https___arxiv_org_abs_2305_19556
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
Park, Se Jin
Kim, Minsu
Choi, Jeongsoo
Ro, Yong Man
Computer Vision and Pattern Recognition
Artificial Intelligence
Sound
Audio and Speech Processing
Image and Video Processing
Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or following phones, the articulation of a phone varies upon the phonetic context. Therefore, modeling lip motion with the phonetic context can generate more spatio-temporally aligned lip movement. In this respect, we investigate the phonetic context in generating lip motion for talking face generation. We propose Context-Aware Lip-Sync framework (CALS), which explicitly leverages phonetic context to generate lip movement of the target face. CALS is comprised of an Audio-to-Lip module and a Lip-to-Face module. The former is pretrained based on masked learning to map each phone to a contextualized lip motion unit. The contextualized lip motion unit then guides the latter in synthesizing a target identity with context-aware lip motion. From extensive experiments, we verify that simply exploiting the phonetic context in the proposed CALS framework effectively enhances spatio-temporal alignment. We also demonstrate the extent to which the phonetic context assists in lip synchronization and find the effective window size for lip generation to be approximately 1.2 seconds.
title Exploring Phonetic Context-Aware Lip-Sync For Talking Face Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Sound
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2305.19556