FSD50K-Solo: Automated Curation of Single-Source Sound Events

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Ningyuan, Yin, Sile, Yang, Li-Chia, Irvin, Bryce, Quan, Xiao, Stamenovic, Marko, Zhang, Shuo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918525828857856
author Yang, Ningyuan
Yin, Sile
Yang, Li-Chia
Irvin, Bryce
Quan, Xiao
Stamenovic, Marko
Zhang, Shuo
author_facet Yang, Ningyuan
Yin, Sile
Yang, Li-Chia
Irvin, Bryce
Quan, Xiao
Stamenovic, Marko
Zhang, Shuo
contents High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13931
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FSD50K-Solo: Automated Curation of Single-Source Sound Events
Yang, Ningyuan
Yin, Sile
Yang, Li-Chia
Irvin, Bryce
Quan, Xiao
Stamenovic, Marko
Zhang, Shuo
Audio and Speech Processing
Sound
High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora.
title FSD50K-Solo: Automated Curation of Single-Source Sound Events
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2605.13931