NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Zongyang, Chandra, Shreeram Suresh, Ulgen, Ismail Rasim, Mahapatra, Aurosweta, Salman, Ali N., Busso, Carlos, Sisman, Berrak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915589709103104
author Du, Zongyang
Chandra, Shreeram Suresh
Ulgen, Ismail Rasim
Mahapatra, Aurosweta
Salman, Ali N.
Busso, Carlos
Sisman, Berrak
author_facet Du, Zongyang
Chandra, Shreeram Suresh
Ulgen, Ismail Rasim
Mahapatra, Aurosweta
Salman, Ali N.
Busso, Carlos
Sisman, Berrak
contents Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive richness of real-life communication. With the rise of large neural networks, several large-scale speech corpora have emerged and been widely adopted across various speech processing tasks. However, the field of voice conversion (VC) still lacks large-scale, expressive, and real-life speech resources suitable for modeling natural prosody and emotion. To fill this gap, we release NaturalVoices (NV), the first large-scale spontaneous podcast dataset specifically designed for emotion-aware voice conversion. It comprises 5,049 hours of spontaneous podcast recordings with automatic annotations for emotion (categorical and attribute-based), speech quality, transcripts, speaker identity, and sound events. The dataset captures expressive emotional variation across thousands of speakers, diverse topics, and natural speaking styles. We also provide an open-source pipeline with modular annotation tools and flexible filtering, enabling researchers to construct customized subsets for a wide range of VC tasks. Experiments demonstrate that NaturalVoices supports the development of robust and generalizable VC models capable of producing natural, expressive speech, while revealing limitations of current architectures when applied to large-scale spontaneous data. These results suggest that NaturalVoices is both a valuable resource and a challenging benchmark for advancing the field of voice conversion. Dataset is available at: https://huggingface.co/JHU-SmileLab
format Preprint
id arxiv_https___arxiv_org_abs_2511_00256
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
Du, Zongyang
Chandra, Shreeram Suresh
Ulgen, Ismail Rasim
Mahapatra, Aurosweta
Salman, Ali N.
Busso, Carlos
Sisman, Berrak
Audio and Speech Processing
Machine Learning
Sound
Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive richness of real-life communication. With the rise of large neural networks, several large-scale speech corpora have emerged and been widely adopted across various speech processing tasks. However, the field of voice conversion (VC) still lacks large-scale, expressive, and real-life speech resources suitable for modeling natural prosody and emotion. To fill this gap, we release NaturalVoices (NV), the first large-scale spontaneous podcast dataset specifically designed for emotion-aware voice conversion. It comprises 5,049 hours of spontaneous podcast recordings with automatic annotations for emotion (categorical and attribute-based), speech quality, transcripts, speaker identity, and sound events. The dataset captures expressive emotional variation across thousands of speakers, diverse topics, and natural speaking styles. We also provide an open-source pipeline with modular annotation tools and flexible filtering, enabling researchers to construct customized subsets for a wide range of VC tasks. Experiments demonstrate that NaturalVoices supports the development of robust and generalizable VC models capable of producing natural, expressive speech, while revealing limitations of current architectures when applied to large-scale spontaneous data. These results suggest that NaturalVoices is both a valuable resource and a challenging benchmark for advancing the field of voice conversion. Dataset is available at: https://huggingface.co/JHU-SmileLab
title NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2511.00256