Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Haorui, Shang, Zengqiang, Wang, Chaoren, Li, Xuyuan, Gu, Yicheng, Hua, Hua, Liu, Liwei, Yang, Chen, Li, Jiaqi, Shi, Peiyang, Wang, Yuancheng, Chen, Kai, Zhang, Pengyuan, Wu, Zhizheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915537340071936
author He, Haorui
Shang, Zengqiang
Wang, Chaoren
Li, Xuyuan
Gu, Yicheng
Hua, Hua
Liu, Liwei
Yang, Chen
Li, Jiaqi
Shi, Peiyang
Wang, Yuancheng
Chen, Kai
Zhang, Pengyuan
Wu, Zhizheng
author_facet He, Haorui
Shang, Zengqiang
Wang, Chaoren
Li, Xuyuan
Gu, Yicheng
Hua, Hua
Liu, Liwei
Yang, Chen
Li, Jiaqi
Shi, Peiyang
Wang, Yuancheng
Chen, Kai
Zhang, Pengyuan
Wu, Zhizheng
contents Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on audio-book datasets limited to formal, read-aloud speaking styles. To address this limitation, we introduce Emilia-Pipe, an open-source preprocessing pipeline designed to extract high-quality training data from valuable yet under-explored in-the-wild sources that capture spontaneous human speech in real-world contexts. Using Emilia-Pipe, we construct Emilia, which comprises over 101k hours of speech across six languages: English, Chinese, German, French, Japanese, and Korean. Furthermore, we expand Emilia to Emilia-Large, a dataset exceeding 216k hours, making it one of the largest open-source speech generation resources available. Extensive experiments show that Emilia-trained models produce markedly more spontaneous, human-like speech than those trained on traditional audio-book datasets, while matching their intelligibility. These models better capture diverse speaker timbres and the full spectrum of real-world conversational styles. Our work also highlights the importance of scaling dataset size for advancing speech generation performance and validates the effectiveness of Emilia for both multilingual and crosslingual speech generation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15907
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
He, Haorui
Shang, Zengqiang
Wang, Chaoren
Li, Xuyuan
Gu, Yicheng
Hua, Hua
Liu, Liwei
Yang, Chen
Li, Jiaqi
Shi, Peiyang
Wang, Yuancheng
Chen, Kai
Zhang, Pengyuan
Wu, Zhizheng
Sound
Computation and Language
Audio and Speech Processing
Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on audio-book datasets limited to formal, read-aloud speaking styles. To address this limitation, we introduce Emilia-Pipe, an open-source preprocessing pipeline designed to extract high-quality training data from valuable yet under-explored in-the-wild sources that capture spontaneous human speech in real-world contexts. Using Emilia-Pipe, we construct Emilia, which comprises over 101k hours of speech across six languages: English, Chinese, German, French, Japanese, and Korean. Furthermore, we expand Emilia to Emilia-Large, a dataset exceeding 216k hours, making it one of the largest open-source speech generation resources available. Extensive experiments show that Emilia-trained models produce markedly more spontaneous, human-like speech than those trained on traditional audio-book datasets, while matching their intelligibility. These models better capture diverse speaker timbres and the full spectrum of real-world conversational styles. Our work also highlights the importance of scaling dataset size for advancing speech generation performance and validates the effectiveness of Emilia for both multilingual and crosslingual speech generation tasks.
title Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2501.15907