Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Han, Wooseok, Kang, Minki, Kim, Changhun, Yang, Eunho
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912171683741696
author Han, Wooseok
Kang, Minki
Kim, Changhun
Yang, Eunho
author_facet Han, Wooseok
Kang, Minki
Kim, Changhun
Yang, Eunho
contents Speaker-adaptive Text-to-Speech (TTS) synthesis has attracted considerable attention due to its broad range of applications, such as personalized voice assistant services. While several approaches have been proposed, they often exhibit high sensitivity to either the quantity or the quality of target speech samples. To address these limitations, we introduce Stable-TTS, a novel speaker-adaptive TTS framework that leverages a small subset of a high-quality pre-training dataset, referred to as prior samples. Specifically, Stable-TTS achieves prosody consistency by leveraging the high-quality prosody of prior samples, while effectively capturing the timbre of the target speaker. Additionally, it employs a prior-preservation loss during fine-tuning to maintain the synthesis ability for prior samples to prevent overfitting on target samples. Extensive experiments demonstrate the effectiveness of Stable-TTS even under limited amounts of and noisy target speech samples.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
Han, Wooseok
Kang, Minki
Kim, Changhun
Yang, Eunho
Sound
Artificial Intelligence
Audio and Speech Processing
Speaker-adaptive Text-to-Speech (TTS) synthesis has attracted considerable attention due to its broad range of applications, such as personalized voice assistant services. While several approaches have been proposed, they often exhibit high sensitivity to either the quantity or the quality of target speech samples. To address these limitations, we introduce Stable-TTS, a novel speaker-adaptive TTS framework that leverages a small subset of a high-quality pre-training dataset, referred to as prior samples. Specifically, Stable-TTS achieves prosody consistency by leveraging the high-quality prosody of prior samples, while effectively capturing the timbre of the target speaker. Additionally, it employs a prior-preservation loss during fine-tuning to maintain the synthesis ability for prior samples to prevent overfitting on target samples. Extensive experiments demonstrate the effectiveness of Stable-TTS even under limited amounts of and noisy target speech samples.
title Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.20155