Scaling Rich Style-Prompted Text-to-Speech Datasets

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Diwan, Anuj, Zheng, Zhisheng, Harwath, David, Choi, Eunsol
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916967803256832
author Diwan, Anuj
Zheng, Zhisheng
Harwath, David
Choi, Eunsol
author_facet Diwan, Anuj
Zheng, Zhisheng
Harwath, David
Choi, Eunsol
contents We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttural, nasal, pained) have been explored in small-scale human-annotated datasets, existing large-scale datasets only cover basic tags (e.g. low-pitched, slow, loud). We combine off-the-shelf text and speech embedders, classifiers and an audio language model to automatically scale rich tag annotations for the first time. ParaSpeechCaps covers a total of 59 style tags, including both speaker-level intrinsic tags and utterance-level situational tags. It consists of 342 hours of human-labelled data (PSC-Base) and 2427 hours of automatically annotated data (PSC-Scaled). We finetune Parler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, and achieve improved style consistency (+7.9% Consistency MOS) and speech quality (+15.5% Naturalness MOS) over the best performing baseline that combines existing rich style tag datasets. We ablate several of our dataset design choices to lay the foundation for future work in this space. Our dataset, models and code are released at https://github.com/ajd12342/paraspeechcaps .
format Preprint
id arxiv_https___arxiv_org_abs_2503_04713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Rich Style-Prompted Text-to-Speech Datasets
Diwan, Anuj
Zheng, Zhisheng
Harwath, David
Choi, Eunsol
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttural, nasal, pained) have been explored in small-scale human-annotated datasets, existing large-scale datasets only cover basic tags (e.g. low-pitched, slow, loud). We combine off-the-shelf text and speech embedders, classifiers and an audio language model to automatically scale rich tag annotations for the first time. ParaSpeechCaps covers a total of 59 style tags, including both speaker-level intrinsic tags and utterance-level situational tags. It consists of 342 hours of human-labelled data (PSC-Base) and 2427 hours of automatically annotated data (PSC-Scaled). We finetune Parler-TTS, an open-source style-prompted TTS model, on ParaSpeechCaps, and achieve improved style consistency (+7.9% Consistency MOS) and speech quality (+15.5% Naturalness MOS) over the best performing baseline that combines existing rich style tag datasets. We ablate several of our dataset design choices to lay the foundation for future work in this space. Our dataset, models and code are released at https://github.com/ajd12342/paraspeechcaps .
title Scaling Rich Style-Prompted Text-to-Speech Datasets
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2503.04713