SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Ziqiao, Hu, Wentao, Shi, Yue, Zhu, Xiangyu, Zhang, Xiaomei, Zhao, Hao, He, Jun, Liu, Hongyan, Fan, Zhaoxin
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917652016922624
author Peng, Ziqiao
Hu, Wentao
Shi, Yue
Zhu, Xiangyu
Zhang, Xiaomei
Zhao, Hao
He, Jun
Liu, Hongyan
Fan, Zhaoxin
author_facet Peng, Ziqiao
Hu, Wentao
Shi, Yue
Zhu, Xiangyu
Zhang, Xiaomei
Zhao, Hao
He, Jun
Liu, Hongyan
Fan, Zhaoxin
contents Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity, while Neural Radiance Fields (NeRF) methods, although they can address this issue, often produce mismatched lip movements, inadequate facial expressions, and unstable head poses. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expressions, and head poses. The absence of these synchronizations is a fundamental flaw, leading to unrealistic and artificial outcomes. To address the critical issue of synchronization, identified as the "devil" in creating realistic talking heads, we introduce SyncTalk. This NeRF-based method effectively maintains subject identity, enhancing synchronization and realism in talking head synthesis. SyncTalk employs a Face-Sync Controller to align lip movements with speech and innovatively uses a 3D facial blendshape model to capture accurate facial expressions. Our Head-Sync Stabilizer optimizes head poses, achieving more natural head movements. The Portrait-Sync Generator restores hair details and blends the generated head with the torso for a seamless visual experience. Extensive experiments and user studies demonstrate that SyncTalk outperforms state-of-the-art methods in synchronization and realism. We recommend watching the supplementary video: https://ziqiaopeng.github.io/synctalk
format Preprint
id arxiv_https___arxiv_org_abs_2311_17590
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
Peng, Ziqiao
Hu, Wentao
Shi, Yue
Zhu, Xiangyu
Zhang, Xiaomei
Zhao, Hao
He, Jun
Liu, Hongyan
Fan, Zhaoxin
Computer Vision and Pattern Recognition
Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity, while Neural Radiance Fields (NeRF) methods, although they can address this issue, often produce mismatched lip movements, inadequate facial expressions, and unstable head poses. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expressions, and head poses. The absence of these synchronizations is a fundamental flaw, leading to unrealistic and artificial outcomes. To address the critical issue of synchronization, identified as the "devil" in creating realistic talking heads, we introduce SyncTalk. This NeRF-based method effectively maintains subject identity, enhancing synchronization and realism in talking head synthesis. SyncTalk employs a Face-Sync Controller to align lip movements with speech and innovatively uses a 3D facial blendshape model to capture accurate facial expressions. Our Head-Sync Stabilizer optimizes head poses, achieving more natural head movements. The Portrait-Sync Generator restores hair details and blends the generated head with the torso for a seamless visual experience. Extensive experiments and user studies demonstrate that SyncTalk outperforms state-of-the-art methods in synchronization and realism. We recommend watching the supplementary video: https://ziqiaopeng.github.io/synctalk
title SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.17590