Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Haorui, Shang, Zengqiang, Wang, Chaoren, Li, Xuyuan, Gu, Yicheng, Hua, Hua, Liu, Liwei, Yang, Chen, Li, Jiaqi, Shi, Peiyang, Wang, Yuancheng, Chen, Kai, Zhang, Pengyuan, Wu, Zhizheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912018586402816
author He, Haorui
Shang, Zengqiang
Wang, Chaoren
Li, Xuyuan
Gu, Yicheng
Hua, Hua
Liu, Liwei
Yang, Chen
Li, Jiaqi
Shi, Peiyang
Wang, Yuancheng
Chen, Kai
Zhang, Pengyuan
Wu, Zhizheng
author_facet He, Haorui
Shang, Zengqiang
Wang, Chaoren
Li, Xuyuan
Gu, Yicheng
Hua, Hua
Liu, Liwei
Yang, Chen
Li, Jiaqi
Shi, Peiyang
Wang, Yuancheng
Chen, Kai
Zhang, Pengyuan
Wu, Zhizheng
contents Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05361
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
He, Haorui
Shang, Zengqiang
Wang, Chaoren
Li, Xuyuan
Gu, Yicheng
Hua, Hua
Liu, Liwei
Yang, Chen
Li, Jiaqi
Shi, Peiyang
Wang, Yuancheng
Chen, Kai
Zhang, Pengyuan
Wu, Zhizheng
Audio and Speech Processing
Computation and Language
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/.
title Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2407.05361