ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Shengkui, Pan, Zexu, Ma, Bin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916809432629248
author Zhao, Shengkui
Pan, Zexu
Ma, Bin
author_facet Zhao, Shengkui
Pan, Zexu
Ma, Bin
contents This paper introduces ClearerVoice-Studio, an open-source, AI-powered speech processing toolkit designed to bridge cutting-edge research and practical application. Unlike broad platforms like SpeechBrain and ESPnet, ClearerVoice-Studio focuses on interconnected speech tasks of speech enhancement, separation, super-resolution, and multimodal target speaker extraction. A key advantage is its state-of-the-art pretrained models, including FRCRN with 3 million uses and MossFormer with 2.5 million uses, optimized for real-world scenarios. It also offers model optimization tools, multi-format audio support, the SpeechScore evaluation toolkit, and user-friendly interfaces, catering to researchers, developers, and end-users. Its rapid adoption attracting 3000 GitHub stars and 239 forks highlights its academic and industrial impact. This paper details ClearerVoice-Studio's capabilities, architectures, training strategies, benchmarks, community impact, and future plan. Source code is available at https://github.com/modelscope/ClearerVoice-Studio.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19398
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
Zhao, Shengkui
Pan, Zexu
Ma, Bin
Sound
Audio and Speech Processing
This paper introduces ClearerVoice-Studio, an open-source, AI-powered speech processing toolkit designed to bridge cutting-edge research and practical application. Unlike broad platforms like SpeechBrain and ESPnet, ClearerVoice-Studio focuses on interconnected speech tasks of speech enhancement, separation, super-resolution, and multimodal target speaker extraction. A key advantage is its state-of-the-art pretrained models, including FRCRN with 3 million uses and MossFormer with 2.5 million uses, optimized for real-world scenarios. It also offers model optimization tools, multi-format audio support, the SpeechScore evaluation toolkit, and user-friendly interfaces, catering to researchers, developers, and end-users. Its rapid adoption attracting 3000 GitHub stars and 239 forks highlights its academic and industrial impact. This paper details ClearerVoice-Studio's capabilities, architectures, training strategies, benchmarks, community impact, and future plan. Source code is available at https://github.com/modelscope/ClearerVoice-Studio.
title ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.19398