MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sung-Bin, Kim, Chae-Yeon, Lee, Son, Gihun, Hyun-Bin, Oh, Ju, Janghoon, Nam, Suekyeong, Oh, Tae-Hyun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914842416250880
author Sung-Bin, Kim
Chae-Yeon, Lee
Son, Gihun
Hyun-Bin, Oh
Ju, Janghoon
Nam, Suekyeong
Oh, Tae-Hyun
author_facet Sung-Bin, Kim
Chae-Yeon, Lee
Son, Gihun
Hyun-Bin, Oh
Ju, Janghoon
Nam, Suekyeong
Oh, Tae-Hyun
contents Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack of datasets covering a broad spectrum of facial movements across languages. In this work, we introduce a novel task to generate 3D talking heads from speeches of diverse languages. We collect a new multilingual 2D video dataset comprising over 420 hours of talking videos in 20 languages. With our proposed dataset, we present a multilingually enhanced model that incorporates language-specific style embeddings, enabling it to capture the unique mouth movements associated with each language. Additionally, we present a metric for assessing lip-sync accuracy in multilingual settings. We demonstrate that training a 3D talking head model with our proposed dataset significantly enhances its multilingual performance. Codes and datasets are available at https://multi-talk.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14272
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
Sung-Bin, Kim
Chae-Yeon, Lee
Son, Gihun
Hyun-Bin, Oh
Ju, Janghoon
Nam, Suekyeong
Oh, Tae-Hyun
Computer Vision and Pattern Recognition
Graphics
Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack of datasets covering a broad spectrum of facial movements across languages. In this work, we introduce a novel task to generate 3D talking heads from speeches of diverse languages. We collect a new multilingual 2D video dataset comprising over 420 hours of talking videos in 20 languages. With our proposed dataset, we present a multilingually enhanced model that incorporates language-specific style embeddings, enabling it to capture the unique mouth movements associated with each language. Additionally, we present a metric for assessing lip-sync accuracy in multilingual settings. We demonstrate that training a 3D talking head model with our proposed dataset significantly enhances its multilingual performance. Codes and datasets are available at https://multi-talk.github.io/.
title MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2406.14272