SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Yicheng, Wang, Chaoren, Zhang, Junan, Zhang, Xueyao, Fang, Zihao, He, Haorui, Wu, Zhizheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918020374331392
author Gu, Yicheng
Wang, Chaoren
Zhang, Junan
Zhang, Xueyao
Fang, Zihao
He, Haorui
Wu, Zhizheng
author_facet Gu, Yicheng
Wang, Chaoren
Zhang, Junan
Zhang, Xueyao
Fang, Zihao
He, Haorui
Wu, Zhizheng
contents The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present SingNet, an extensive, diverse, and in-the-wild singing voice dataset. Specifically, we propose a data processing pipeline to extract ready-to-use training data from sample packs and songs on the internet, forming 3000 hours of singing voices in various languages and styles. Furthermore, to facilitate the use and demonstrate the effectiveness of SingNet, we pre-train and open-source various state-of-the-art (SOTA) models on Wav2vec2, BigVGAN, and NSF-HiFiGAN based on our collected singing voice data. We also conduct benchmark experiments on Automatic Lyric Transcription (ALT), Neural Vocoder, and Singing Voice Conversion (SVC). Audio demos are available at: https://singnet-dataset.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
Gu, Yicheng
Wang, Chaoren
Zhang, Junan
Zhang, Xueyao
Fang, Zihao
He, Haorui
Wu, Zhizheng
Sound
Audio and Speech Processing
The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present SingNet, an extensive, diverse, and in-the-wild singing voice dataset. Specifically, we propose a data processing pipeline to extract ready-to-use training data from sample packs and songs on the internet, forming 3000 hours of singing voices in various languages and styles. Furthermore, to facilitate the use and demonstrate the effectiveness of SingNet, we pre-train and open-source various state-of-the-art (SOTA) models on Wav2vec2, BigVGAN, and NSF-HiFiGAN based on our collected singing voice data. We also conduct benchmark experiments on Automatic Lyric Transcription (ALT), Neural Vocoder, and Singing Voice Conversion (SVC). Audio demos are available at: https://singnet-dataset.github.io/.
title SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.09325