Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Marques, Leonardo B. de M. M., Ueda, Lucas H., Neto, Mário U., Simões, Flávio O., Runstein, Fernando, Bó, Bianca Dal, Costa, Paula D. P.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913536009043968
author Marques, Leonardo B. de M. M.
Ueda, Lucas H.
Neto, Mário U.
Simões, Flávio O.
Runstein, Fernando
Bó, Bianca Dal
Costa, Paula D. P.
author_facet Marques, Leonardo B. de M. M.
Ueda, Lucas H.
Neto, Mário U.
Simões, Flávio O.
Runstein, Fernando
Bó, Bianca Dal
Costa, Paula D. P.
contents The goal of cross-speaker style transfer in TTS is to transfer a speech style from a source speaker with expressive data to a target speaker with only neutral data. In this context, we propose using a pre-trained singing voice conversion (SVC) model to convert the expressive data into the target speaker's voice. In the conversion process, we apply a fundamental frequency (F0) matching technique to mitigate tonal variances between speakers with significant timbral differences. A style classifier filter is proposed to select the most expressive output audios for the TTS training. Our approach is comparable to state-of-the-art with only a few minutes of neutral data from the target speaker, while other methods require hours. A perceptual assessment showed improvements brought by the SVC and the style filter in naturalness and style intensity for the styles that display more vocal effort. Also, increased speaker similarity is obtained with the proposed F0 matching algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05620
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
Marques, Leonardo B. de M. M.
Ueda, Lucas H.
Neto, Mário U.
Simões, Flávio O.
Runstein, Fernando
Bó, Bianca Dal
Costa, Paula D. P.
Audio and Speech Processing
The goal of cross-speaker style transfer in TTS is to transfer a speech style from a source speaker with expressive data to a target speaker with only neutral data. In this context, we propose using a pre-trained singing voice conversion (SVC) model to convert the expressive data into the target speaker's voice. In the conversion process, we apply a fundamental frequency (F0) matching technique to mitigate tonal variances between speakers with significant timbral differences. A style classifier filter is proposed to select the most expressive output audios for the TTS training. Our approach is comparable to state-of-the-art with only a few minutes of neutral data from the target speaker, while other methods require hours. A perceptual assessment showed improvements brought by the SVC and the style filter in naturalness and style intensity for the styles that display more vocal effort. Also, increased speaker similarity is obtained with the proposed F0 matching algorithm.
title Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
topic Audio and Speech Processing
url https://arxiv.org/abs/2410.05620