CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Helin, Hai, Jiarui, Chong, Dading, Thakkar, Karan, Feng, Tiantian, Yang, Dongchao, Lee, Junhyeok, Thebaud, Thomas, Velazquez, Laureano Moro, Villalba, Jesus, Qin, Zengyi, Narayanan, Shrikanth, Elhiali, Mounya, Dehak, Najim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912607210831872
author Wang, Helin
Hai, Jiarui
Chong, Dading
Thakkar, Karan
Feng, Tiantian
Yang, Dongchao
Lee, Junhyeok
Thebaud, Thomas
Velazquez, Laureano Moro
Villalba, Jesus
Qin, Zengyi
Narayanan, Shrikanth
Elhiali, Mounya
Dehak, Najim
author_facet Wang, Helin
Hai, Jiarui
Chong, Dading
Thakkar, Karan
Feng, Tiantian
Yang, Dongchao
Lee, Junhyeok
Thebaud, Thomas
Velazquez, Laureano Moro
Villalba, Jesus
Qin, Zengyi
Narayanan, Shrikanth
Elhiali, Mounya
Dehak, Najim
contents Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack of standardized, comprehensive datasets and limited research on downstream tasks built upon CapTTS. To address these gaps, we introduce CapSpeech, a new benchmark designed for a series of CapTTS-related tasks, including style-captioned text-to-speech synthesis with sound events (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS), and text-to-speech synthesis for chat agent (AgentTTS). CapSpeech comprises over 10 million machine-annotated audio-caption pairs and nearly 0.36 million human-annotated audio-caption pairs. In addition, we introduce two new datasets collected and recorded by a professional voice actor and experienced audio engineers, specifically for the AgentTTS and CapTTS-SE tasks. Alongside the datasets, we conduct comprehensive experiments using both autoregressive and non-autoregressive models on CapSpeech. Our results demonstrate high-fidelity and highly intelligible speech synthesis across a diverse range of speaking styles. To the best of our knowledge, CapSpeech is the largest available dataset offering comprehensive annotations for CapTTS-related tasks. The experiments and findings further provide valuable insights into the challenges of developing CapTTS systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Wang, Helin
Hai, Jiarui
Chong, Dading
Thakkar, Karan
Feng, Tiantian
Yang, Dongchao
Lee, Junhyeok
Thebaud, Thomas
Velazquez, Laureano Moro
Villalba, Jesus
Qin, Zengyi
Narayanan, Shrikanth
Elhiali, Mounya
Dehak, Najim
Audio and Speech Processing
Artificial Intelligence
Sound
Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack of standardized, comprehensive datasets and limited research on downstream tasks built upon CapTTS. To address these gaps, we introduce CapSpeech, a new benchmark designed for a series of CapTTS-related tasks, including style-captioned text-to-speech synthesis with sound events (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS), and text-to-speech synthesis for chat agent (AgentTTS). CapSpeech comprises over 10 million machine-annotated audio-caption pairs and nearly 0.36 million human-annotated audio-caption pairs. In addition, we introduce two new datasets collected and recorded by a professional voice actor and experienced audio engineers, specifically for the AgentTTS and CapTTS-SE tasks. Alongside the datasets, we conduct comprehensive experiments using both autoregressive and non-autoregressive models on CapSpeech. Our results demonstrate high-fidelity and highly intelligible speech synthesis across a diverse range of speaking styles. To the best of our knowledge, CapSpeech is the largest available dataset offering comprehensive annotations for CapTTS-related tasks. The experiments and findings further provide valuable insights into the challenges of developing CapTTS systems.
title CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2506.02863