Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Minsu, Ma, Pingchuan, Chen, Honglie, Petridis, Stavros, Pantic, Maja
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910967437197312
author Kim, Minsu
Ma, Pingchuan
Chen, Honglie
Petridis, Stavros
Pantic, Maja
author_facet Kim, Minsu
Ma, Pingchuan
Chen, Honglie
Petridis, Stavros
Pantic, Maja
contents This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with natural text description. Specifically, we aim to mitigate the following three challenges in face-driven TTS systems. 1) To overcome the limited audio quality of audio-visual speech corpora, we propose a training method that additionally utilizes high-quality audio-only speech corpora. 2) To generate voices not only from real human faces but also from artistic portraits, we propose augmenting the input face image with stylization. 3) To consider one-to-many possibilities in face-to-voice mapping and ensure consistent voice generation at the same time, we propose to first employ sampling-based decoding and then use prompting with generated speech samples. Experimental results validate the proposed model's effectiveness in face-driven voice synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18972
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
Kim, Minsu
Ma, Pingchuan
Chen, Honglie
Petridis, Stavros
Pantic, Maja
Audio and Speech Processing
Artificial Intelligence
This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with natural text description. Specifically, we aim to mitigate the following three challenges in face-driven TTS systems. 1) To overcome the limited audio quality of audio-visual speech corpora, we propose a training method that additionally utilizes high-quality audio-only speech corpora. 2) To generate voices not only from real human faces but also from artistic portraits, we propose augmenting the input face image with stylization. 3) To consider one-to-many possibilities in face-to-voice mapping and ensure consistent voice generation at the same time, we propose to first employ sampling-based decoding and then use prompting with generated speech samples. Experimental results validate the proposed model's effectiveness in face-driven voice synthesis.
title Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2505.18972