FSD50K-Solo: Automated Curation of Single-Source Sound Events
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918525828857856 |
|---|---|
| author | Yang, Ningyuan Yin, Sile Yang, Li-Chia Irvin, Bryce Quan, Xiao Stamenovic, Marko Zhang, Shuo |
| author_facet | Yang, Ningyuan Yin, Sile Yang, Li-Chia Irvin, Bryce Quan, Xiao Stamenovic, Marko Zhang, Shuo |
| contents | High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_13931 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FSD50K-Solo: Automated Curation of Single-Source Sound Events Yang, Ningyuan Yin, Sile Yang, Li-Chia Irvin, Bryce Quan, Xiao Stamenovic, Marko Zhang, Shuo Audio and Speech Processing Sound High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound event dataset. The FSD50K dataset, despite being relatively large and open, contains a considerable fraction of multi-source samples where background interference or overlapping events could limit the usefulness of the data. To address this challenge, we introduce a data curation framework designed for large-scale open audio corpora. Our approach leverages a generative diffusion model to synthesize clean single-class events to construct controlled noisy mixtures for supervision. We subsequently employ a pre-trained audio encoder coupled with a discriminative classifier to automatically identify and filter out multi-source samples. Experiments show that our framework achieves strong performance on a human expert-curated test set. Finally, we release FSD50K-Solo, a model-curated subset of FSD50K containing single-source audio samples identified by our method. Beyond FSD50K, our method establishes a scalable paradigm for curating open source audio corpora. |
| title | FSD50K-Solo: Automated Curation of Single-Source Sound Events |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2605.13931 |