Less is More: Data Curation Matters in Scaling Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenda, Zhang, Wangyou, Wang, Wei, Scheibler, Robin, Saijo, Kohei, Cornell, Samuele, Fu, Yihui, Sach, Marvin, Ni, Zhaoheng, Kumar, Anurag, Fingscheidt, Tim, Watanabe, Shinji, Qian, Yanmin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912543193169920
author Li, Chenda
Zhang, Wangyou
Wang, Wei
Scheibler, Robin
Saijo, Kohei
Cornell, Samuele
Fu, Yihui
Sach, Marvin
Ni, Zhaoheng
Kumar, Anurag
Fingscheidt, Tim
Watanabe, Shinji
Qian, Yanmin
author_facet Li, Chenda
Zhang, Wangyou
Wang, Wei
Scheibler, Robin
Saijo, Kohei
Cornell, Samuele
Fu, Yihui
Sach, Marvin
Ni, Zhaoheng
Kumar, Anurag
Fingscheidt, Tim
Watanabe, Shinji
Qian, Yanmin
contents The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Less is More: Data Curation Matters in Scaling Speech Enhancement
Li, Chenda
Zhang, Wangyou
Wang, Wei
Scheibler, Robin
Saijo, Kohei
Cornell, Samuele
Fu, Yihui
Sach, Marvin
Ni, Zhaoheng
Kumar, Anurag
Fingscheidt, Tim
Watanabe, Shinji
Qian, Yanmin
Audio and Speech Processing
Sound
The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
title Less is More: Data Curation Matters in Scaling Speech Enhancement
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.23859