Building speech corpus with diverse voice characteristics for its prompt-based representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Watanabe, Aya, Takamichi, Shinnosuke, Saito, Yuki, Nakata, Wataru, Xin, Detai, Saruwatari, Hiroshi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910375132266496
author Watanabe, Aya
Takamichi, Shinnosuke
Saito, Yuki
Nakata, Wataru
Xin, Detai
Saruwatari, Hiroshi
author_facet Watanabe, Aya
Takamichi, Shinnosuke
Saito, Yuki
Nakata, Wataru
Xin, Detai
Saruwatari, Hiroshi
contents In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice characteristics. While previous research has explored the prompt-based manipulation of voice characteristics, most studies have used pre-recorded speech, which limits the diversity of voice characteristics available. Thus, we aim to address this gap by creating a novel corpus and developing a model for prompt-based manipulation of voice characteristics in text-to-speech synthesis, facilitating a broader range of voice characteristics. Specifically, we propose a method to build a sizable corpus pairing voice characteristics descriptions with corresponding speech samples. This involves automatically gathering voice-related speech data from the Internet, ensuring its quality, and manually annotating it using crowdsourcing. We implement this method with Japanese language data and analyze the results to validate its effectiveness. Subsequently, we propose a construction method of the model to retrieve speech from voice characteristics descriptions based on a contrastive learning method. We train the model using not only conservative contrastive learning but also feature prediction learning to predict quantitative speech features corresponding to voice characteristics. We evaluate the model performance via experiments with the corpus we constructed above.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13353
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Building speech corpus with diverse voice characteristics for its prompt-based representation
Watanabe, Aya
Takamichi, Shinnosuke
Saito, Yuki
Nakata, Wataru
Xin, Detai
Saruwatari, Hiroshi
Sound
Audio and Speech Processing
In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice characteristics. While previous research has explored the prompt-based manipulation of voice characteristics, most studies have used pre-recorded speech, which limits the diversity of voice characteristics available. Thus, we aim to address this gap by creating a novel corpus and developing a model for prompt-based manipulation of voice characteristics in text-to-speech synthesis, facilitating a broader range of voice characteristics. Specifically, we propose a method to build a sizable corpus pairing voice characteristics descriptions with corresponding speech samples. This involves automatically gathering voice-related speech data from the Internet, ensuring its quality, and manually annotating it using crowdsourcing. We implement this method with Japanese language data and analyze the results to validate its effectiveness. Subsequently, we propose a construction method of the model to retrieve speech from voice characteristics descriptions based on a contrastive learning method. We train the model using not only conservative contrastive learning but also feature prediction learning to predict quantitative speech features corresponding to voice characteristics. We evaluate the model performance via experiments with the corpus we constructed above.
title Building speech corpus with diverse voice characteristics for its prompt-based representation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2403.13353