Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910967437197312 |
|---|---|
| author | Kim, Minsu Ma, Pingchuan Chen, Honglie Petridis, Stavros Pantic, Maja |
| author_facet | Kim, Minsu Ma, Pingchuan Chen, Honglie Petridis, Stavros Pantic, Maja |
| contents | This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with natural text description. Specifically, we aim to mitigate the following three challenges in face-driven TTS systems. 1) To overcome the limited audio quality of audio-visual speech corpora, we propose a training method that additionally utilizes high-quality audio-only speech corpora. 2) To generate voices not only from real human faces but also from artistic portraits, we propose augmenting the input face image with stylization. 3) To consider one-to-many possibilities in face-to-voice mapping and ensure consistent voice generation at the same time, we propose to first employ sampling-based decoding and then use prompting with generated speech samples. Experimental results validate the proposed model's effectiveness in face-driven voice synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18972 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis Kim, Minsu Ma, Pingchuan Chen, Honglie Petridis, Stavros Pantic, Maja Audio and Speech Processing Artificial Intelligence This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with natural text description. Specifically, we aim to mitigate the following three challenges in face-driven TTS systems. 1) To overcome the limited audio quality of audio-visual speech corpora, we propose a training method that additionally utilizes high-quality audio-only speech corpora. 2) To generate voices not only from real human faces but also from artistic portraits, we propose augmenting the input face image with stylization. 3) To consider one-to-many possibilities in face-to-voice mapping and ensure consistent voice generation at the same time, we propose to first employ sampling-based decoding and then use prompting with generated speech samples. Experimental results validate the proposed model's effectiveness in face-driven voice synthesis. |
| title | Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis |
| topic | Audio and Speech Processing Artificial Intelligence |
| url | https://arxiv.org/abs/2505.18972 |