Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912018586402816 |
|---|---|
| author | He, Haorui Shang, Zengqiang Wang, Chaoren Li, Xuyuan Gu, Yicheng Hua, Hua Liu, Liwei Yang, Chen Li, Jiaqi Shi, Peiyang Wang, Yuancheng Chen, Kai Zhang, Pengyuan Wu, Zhizheng |
| author_facet | He, Haorui Shang, Zengqiang Wang, Chaoren Li, Xuyuan Gu, Yicheng Hua, Hua Liu, Liwei Yang, Chen Li, Jiaqi Shi, Peiyang Wang, Yuancheng Chen, Kai Zhang, Pengyuan Wu, Zhizheng |
| contents | Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_05361 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation He, Haorui Shang, Zengqiang Wang, Chaoren Li, Xuyuan Gu, Yicheng Hua, Hua Liu, Liwei Yang, Chen Li, Jiaqi Shi, Peiyang Wang, Yuancheng Chen, Kai Zhang, Pengyuan Wu, Zhizheng Audio and Speech Processing Computation and Language Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/. |
| title | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation |
| topic | Audio and Speech Processing Computation and Language |
| url | https://arxiv.org/abs/2407.05361 |