Overview of the Amphion Toolkit (v0.2)
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912228034215936 |
|---|---|
| author | Li, Jiaqi Zhang, Xueyao Wang, Yuancheng He, Haorui Wang, Chaoren Wang, Li Liao, Huan Ao, Junyi Xie, Zeyu Huang, Yiqiao Zhang, Junan Wu, Zhizheng |
| author_facet | Li, Jiaqi Zhang, Xueyao Wang, Yuancheng He, Haorui Wang, Chaoren Wang, Li Liao, Huan Ao, Junyi Xie, Zeyu Huang, Yiqiao Zhang, Junan Wu, Zhizheng |
| contents | Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a versatile framework that supports a variety of generation tasks and models. In this report, we introduce Amphion v0.2, the second major release developed in 2024. This release features a 100K-hour open-source multilingual dataset, a robust data preparation pipeline, and novel models for tasks such as text-to-speech, audio coding, and voice conversion. Furthermore, the report includes multiple tutorials that guide users through the functionalities and usage of the newly released models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_15442 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Overview of the Amphion Toolkit (v0.2) Li, Jiaqi Zhang, Xueyao Wang, Yuancheng He, Haorui Wang, Chaoren Wang, Li Liao, Huan Ao, Junyi Xie, Zeyu Huang, Yiqiao Zhang, Junan Wu, Zhizheng Sound Artificial Intelligence Audio and Speech Processing Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a versatile framework that supports a variety of generation tasks and models. In this report, we introduce Amphion v0.2, the second major release developed in 2024. This release features a 100K-hour open-source multilingual dataset, a robust data preparation pipeline, and novel models for tasks such as text-to-speech, audio coding, and voice conversion. Furthermore, the report includes multiple tutorials that guide users through the functionalities and usage of the newly released models. |
| title | Overview of the Amphion Toolkit (v0.2) |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2501.15442 |