Less is More: Data Curation Matters in Scaling Speech Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912543193169920 |
|---|---|
| author | Li, Chenda Zhang, Wangyou Wang, Wei Scheibler, Robin Saijo, Kohei Cornell, Samuele Fu, Yihui Sach, Marvin Ni, Zhaoheng Kumar, Anurag Fingscheidt, Tim Watanabe, Shinji Qian, Yanmin |
| author_facet | Li, Chenda Zhang, Wangyou Wang, Wei Scheibler, Robin Saijo, Kohei Cornell, Samuele Fu, Yihui Sach, Marvin Ni, Zhaoheng Kumar, Anurag Fingscheidt, Tim Watanabe, Shinji Qian, Yanmin |
| contents | The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_23859 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Less is More: Data Curation Matters in Scaling Speech Enhancement Li, Chenda Zhang, Wangyou Wang, Wei Scheibler, Robin Saijo, Kohei Cornell, Samuele Fu, Yihui Sach, Marvin Ni, Zhaoheng Kumar, Anurag Fingscheidt, Tim Watanabe, Shinji Qian, Yanmin Audio and Speech Processing Sound The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively. |
| title | Less is More: Data Curation Matters in Scaling Speech Enhancement |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2506.23859 |