WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dai, Yuhang, Zhang, Ziyu, Wang, Shuai, Li, Longhao, Guo, Zhao, Zuo, Tianlun, Wang, Shuiyuan, Xue, Hongfei, Wang, Chengyou, Wang, Qing, Xu, Xin, Bu, Hui, Li, Jie, Kang, Jian, Zhang, Binbin, Xie, Lei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908552686206976
author Dai, Yuhang
Zhang, Ziyu
Wang, Shuai
Li, Longhao
Guo, Zhao
Zuo, Tianlun
Wang, Shuiyuan
Xue, Hongfei
Wang, Chengyou
Wang, Qing
Xu, Xin
Bu, Hui
Li, Jie
Kang, Jian
Zhang, Binbin
Xie, Lei
author_facet Dai, Yuhang
Zhang, Ziyu
Wang, Shuai
Li, Longhao
Guo, Zhao
Zuo, Tianlun
Wang, Shuiyuan
Xue, Hongfei
Wang, Chengyou
Wang, Qing
Xu, Xin
Bu, Hui
Li, Jie
Kang, Jian
Zhang, Binbin
Xie, Lei
contents The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constructed using our novel Chuan-Pipeline, a complete data processing framework for dialectal speech. To facilitate rigorous evaluation and demonstrate the corpus's effectiveness, we also release high-quality ASR and TTS benchmarks, WenetSpeech-Chuan-Eval, with manually verified transcriptions. Experiments show that models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and demonstrate results comparable to commercial services. As the largest open-source corpus for Sichuanese dialects, WenetSpeech-Chuan not only lowers the barrier to research in dialectal speech processing but also plays a crucial role in promoting AI equity and mitigating bias in speech technologies. The corpus, benchmarks, models, and receipts are publicly available on our project page.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
Dai, Yuhang
Zhang, Ziyu
Wang, Shuai
Li, Longhao
Guo, Zhao
Zuo, Tianlun
Wang, Shuiyuan
Xue, Hongfei
Wang, Chengyou
Wang, Qing
Xu, Xin
Bu, Hui
Li, Jie
Kang, Jian
Zhang, Binbin
Xie, Lei
Computation and Language
Sound
The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constructed using our novel Chuan-Pipeline, a complete data processing framework for dialectal speech. To facilitate rigorous evaluation and demonstrate the corpus's effectiveness, we also release high-quality ASR and TTS benchmarks, WenetSpeech-Chuan-Eval, with manually verified transcriptions. Experiments show that models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and demonstrate results comparable to commercial services. As the largest open-source corpus for Sichuanese dialects, WenetSpeech-Chuan not only lowers the barrier to research in dialectal speech processing but also plays a crucial role in promoting AI equity and mitigating bias in speech technologies. The corpus, benchmarks, models, and receipts are publicly available on our project page.
title WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
topic Computation and Language
Sound
url https://arxiv.org/abs/2509.18004