VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916663333486592 |
|---|---|
| author | Shi, Jiatong Shim, Hye-jin Tian, Jinchuan Arora, Siddhant Wu, Haibin Petermann, Darius Yip, Jia Qi Zhang, You Tang, Yuxun Zhang, Wangyou Alharthi, Dareen Safar Huang, Yichen Saito, Koichi Han, Jionghao Zhao, Yiwen Donahue, Chris Watanabe, Shinji |
| author_facet | Shi, Jiatong Shim, Hye-jin Tian, Jinchuan Arora, Siddhant Wu, Haibin Petermann, Darius Yip, Jia Qi Zhang, You Tang, Yuxun Zhang, Wangyou Alharthi, Dareen Safar Huang, Yichen Saito, Koichi Han, Jionghao Zhao, Yiwen Donahue, Chris Watanabe, Shinji |
| contents | In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA offers 65 metrics with 729 metric variations based on different configurations. These metrics encompass evaluations utilizing diverse external resources, including matching and non-matching reference audio, text transcriptions, and text captions. As a lightweight yet comprehensive toolkit, VERSA is versatile to support the evaluation of a wide range of downstream scenarios. To demonstrate its capabilities, this work highlights example use cases for VERSA, including audio coding, speech synthesis, speech enhancement, singing synthesis, and music generation. The toolkit is available at https://github.com/wavlab-speech/versa. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_17667 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music Shi, Jiatong Shim, Hye-jin Tian, Jinchuan Arora, Siddhant Wu, Haibin Petermann, Darius Yip, Jia Qi Zhang, You Tang, Yuxun Zhang, Wangyou Alharthi, Dareen Safar Huang, Yichen Saito, Koichi Han, Jionghao Zhao, Yiwen Donahue, Chris Watanabe, Shinji Sound Multimedia Audio and Speech Processing In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA offers 65 metrics with 729 metric variations based on different configurations. These metrics encompass evaluations utilizing diverse external resources, including matching and non-matching reference audio, text transcriptions, and text captions. As a lightweight yet comprehensive toolkit, VERSA is versatile to support the evaluation of a wide range of downstream scenarios. To demonstrate its capabilities, this work highlights example use cases for VERSA, including audio coding, speech synthesis, speech enhancement, singing synthesis, and music generation. The toolkit is available at https://github.com/wavlab-speech/versa. |
| title | VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music |
| topic | Sound Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.17667 |