MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914842416250880 |
|---|---|
| author | Sung-Bin, Kim Chae-Yeon, Lee Son, Gihun Hyun-Bin, Oh Ju, Janghoon Nam, Suekyeong Oh, Tae-Hyun |
| author_facet | Sung-Bin, Kim Chae-Yeon, Lee Son, Gihun Hyun-Bin, Oh Ju, Janghoon Nam, Suekyeong Oh, Tae-Hyun |
| contents | Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack of datasets covering a broad spectrum of facial movements across languages. In this work, we introduce a novel task to generate 3D talking heads from speeches of diverse languages. We collect a new multilingual 2D video dataset comprising over 420 hours of talking videos in 20 languages. With our proposed dataset, we present a multilingually enhanced model that incorporates language-specific style embeddings, enabling it to capture the unique mouth movements associated with each language. Additionally, we present a metric for assessing lip-sync accuracy in multilingual settings. We demonstrate that training a 3D talking head model with our proposed dataset significantly enhances its multilingual performance. Codes and datasets are available at https://multi-talk.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_14272 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset Sung-Bin, Kim Chae-Yeon, Lee Son, Gihun Hyun-Bin, Oh Ju, Janghoon Nam, Suekyeong Oh, Tae-Hyun Computer Vision and Pattern Recognition Graphics Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack of datasets covering a broad spectrum of facial movements across languages. In this work, we introduce a novel task to generate 3D talking heads from speeches of diverse languages. We collect a new multilingual 2D video dataset comprising over 420 hours of talking videos in 20 languages. With our proposed dataset, we present a multilingually enhanced model that incorporates language-specific style embeddings, enabling it to capture the unique mouth movements associated with each language. Additionally, we present a metric for assessing lip-sync accuracy in multilingual settings. We demonstrate that training a 3D talking head model with our proposed dataset significantly enhances its multilingual performance. Codes and datasets are available at https://multi-talk.github.io/. |
| title | MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset |
| topic | Computer Vision and Pattern Recognition Graphics |
| url | https://arxiv.org/abs/2406.14272 |