SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Zeren, Qin, Haibo, Huang, Jiayu, Li, Yixin, Lin, Hui, Duan, Yitao, Ma, Jinwen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911871444975616
author Zhang, Zeren
Qin, Haibo
Huang, Jiayu
Li, Yixin
Lin, Hui
Duan, Yitao
Ma, Jinwen
author_facet Zhang, Zeren
Qin, Haibo
Huang, Jiayu
Li, Yixin
Lin, Hui
Duan, Yitao
Ma, Jinwen
contents Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between tasks and reduce video clarity because the interaction space is limited to the low-level semantic RGB space. To address this issue, we propose an innovative unified framework, SwapTalk, which accomplishes both face swapping and lip synchronization tasks in the same latent space. Referring to recent work on face generation, we choose the VQ-embedding space due to its excellent editability and fidelity performance. To enhance the framework's generalization capabilities for unseen identities, we incorporate identity loss during the training of the face swapping module. Additionally, we introduce expert discriminator supervision within the latent space during the training of the lip synchronization module to elevate synchronization quality. In the evaluation phase, previous studies primarily focused on the self-reconstruction of lip movements in synchronous audio-visual videos. To better approximate real-world applications, we expand the evaluation scope to asynchronous audio-video scenarios. Furthermore, we introduce a novel identity consistency metric to more comprehensively assess the identity consistency over time series in generated facial videos. Experimental results on the HDTF demonstrate that our method significantly surpasses existing techniques in video quality, lip synchronization accuracy, face swapping fidelity, and identity consistency. Our demo is available at http://swaptalk.cc.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05636
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
Zhang, Zeren
Qin, Haibo
Huang, Jiayu
Li, Yixin
Lin, Hui
Duan, Yitao
Ma, Jinwen
Computer Vision and Pattern Recognition
Artificial Intelligence
Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between tasks and reduce video clarity because the interaction space is limited to the low-level semantic RGB space. To address this issue, we propose an innovative unified framework, SwapTalk, which accomplishes both face swapping and lip synchronization tasks in the same latent space. Referring to recent work on face generation, we choose the VQ-embedding space due to its excellent editability and fidelity performance. To enhance the framework's generalization capabilities for unseen identities, we incorporate identity loss during the training of the face swapping module. Additionally, we introduce expert discriminator supervision within the latent space during the training of the lip synchronization module to elevate synchronization quality. In the evaluation phase, previous studies primarily focused on the self-reconstruction of lip movements in synchronous audio-visual videos. To better approximate real-world applications, we expand the evaluation scope to asynchronous audio-video scenarios. Furthermore, we introduce a novel identity consistency metric to more comprehensively assess the identity consistency over time series in generated facial videos. Experimental results on the HDTF demonstrate that our method significantly surpasses existing techniques in video quality, lip synchronization accuracy, face swapping fidelity, and identity consistency. Our demo is available at http://swaptalk.cc.
title SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.05636