Overview of the Amphion Toolkit (v0.2)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiaqi, Zhang, Xueyao, Wang, Yuancheng, He, Haorui, Wang, Chaoren, Wang, Li, Liao, Huan, Ao, Junyi, Xie, Zeyu, Huang, Yiqiao, Zhang, Junan, Wu, Zhizheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912228034215936
author Li, Jiaqi
Zhang, Xueyao
Wang, Yuancheng
He, Haorui
Wang, Chaoren
Wang, Li
Liao, Huan
Ao, Junyi
Xie, Zeyu
Huang, Yiqiao
Zhang, Junan
Wu, Zhizheng
author_facet Li, Jiaqi
Zhang, Xueyao
Wang, Yuancheng
He, Haorui
Wang, Chaoren
Wang, Li
Liao, Huan
Ao, Junyi
Xie, Zeyu
Huang, Yiqiao
Zhang, Junan
Wu, Zhizheng
contents Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a versatile framework that supports a variety of generation tasks and models. In this report, we introduce Amphion v0.2, the second major release developed in 2024. This release features a 100K-hour open-source multilingual dataset, a robust data preparation pipeline, and novel models for tasks such as text-to-speech, audio coding, and voice conversion. Furthermore, the report includes multiple tutorials that guide users through the functionalities and usage of the newly released models.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15442
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Overview of the Amphion Toolkit (v0.2)
Li, Jiaqi
Zhang, Xueyao
Wang, Yuancheng
He, Haorui
Wang, Chaoren
Wang, Li
Liao, Huan
Ao, Junyi
Xie, Zeyu
Huang, Yiqiao
Zhang, Junan
Wu, Zhizheng
Sound
Artificial Intelligence
Audio and Speech Processing
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a versatile framework that supports a variety of generation tasks and models. In this report, we introduce Amphion v0.2, the second major release developed in 2024. This release features a 100K-hour open-source multilingual dataset, a robust data preparation pipeline, and novel models for tasks such as text-to-speech, audio coding, and voice conversion. Furthermore, the report includes multiple tutorials that guide users through the functionalities and usage of the newly released models.
title Overview of the Amphion Toolkit (v0.2)
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2501.15442