Amphion: An Open-Source Audio, Music and Speech Generation Toolkit

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xueyao, Xue, Liumeng, Gu, Yicheng, Wang, Yuancheng, Li, Jiaqi, He, Haorui, Wang, Chaoren, Liu, Songting, Chen, Xi, Zhang, Junan, Fang, Zihao, Chen, Haopeng, Tang, Tze Ying, Zou, Lexiao, Wang, Mingxuan, Han, Jun, Chen, Kai, Li, Haizhou, Wu, Zhizheng
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916395191631872
author Zhang, Xueyao
Xue, Liumeng
Gu, Yicheng
Wang, Yuancheng
Li, Jiaqi
He, Haorui
Wang, Chaoren
Liu, Songting
Chen, Xi
Zhang, Junan
Fang, Zihao
Chen, Haopeng
Tang, Tze Ying
Zou, Lexiao
Wang, Mingxuan
Han, Jun
Chen, Kai
Li, Haizhou
Wu, Zhizheng
author_facet Zhang, Xueyao
Xue, Liumeng
Gu, Yicheng
Wang, Yuancheng
Li, Jiaqi
He, Haorui
Wang, Chaoren
Liu, Songting
Chen, Xi
Zhang, Junan
Fang, Zihao
Chen, Haopeng
Tang, Tze Ying
Zou, Lexiao
Wang, Mingxuan
Han, Jun
Chen, Kai
Li, Haizhou
Wu, Zhizheng
contents Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.
format Preprint
id arxiv_https___arxiv_org_abs_2312_09911
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
Zhang, Xueyao
Xue, Liumeng
Gu, Yicheng
Wang, Yuancheng
Li, Jiaqi
He, Haorui
Wang, Chaoren
Liu, Songting
Chen, Xi
Zhang, Junan
Fang, Zihao
Chen, Haopeng
Tang, Tze Ying
Zou, Lexiao
Wang, Mingxuan
Han, Jun
Chen, Kai
Li, Haizhou
Wu, Zhizheng
Sound
Audio and Speech Processing
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.
title Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2312.09911