Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914072703795200 |
|---|---|
| author | Di Bernardo, Matías Misley, Emmanuel Correa, Ignacio Iacovelli, Mateo García Mellino, Simón Barrios, Gala Lucía Gonzalez |
| author_facet | Di Bernardo, Matías Misley, Emmanuel Correa, Ignacio Iacovelli, Mateo García Mellino, Simón Barrios, Gala Lucía Gonzalez |
| contents | This work introduces a reproducible, metric-driven methodology to evaluate preprocessing pipelines for in-the-wild TTS corpora generation. We apply a custom low-cost pipeline to the first in-the-wild Argentine Spanish collection and compare 24 pipeline configurations combining different denoising and quality filtering variants. Evaluation relies on complementary objective measures (PESQ, SI-SDR, SNR), acoustic descriptors (T30, C50), and speech-preservation metrics (F0-STD, MCD). Results expose trade-offs between dataset size, signal quality, and voice preservation; where denoising variants with permissive filtering provide the best overall compromise for our testbed. The proposed methodology allows selecting pipeline configurations without training TTS models for each subset, accelerating and reducing the cost of preprocessing development for low-resource settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03111 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets Di Bernardo, Matías Misley, Emmanuel Correa, Ignacio Iacovelli, Mateo García Mellino, Simón Barrios, Gala Lucía Gonzalez Audio and Speech Processing This work introduces a reproducible, metric-driven methodology to evaluate preprocessing pipelines for in-the-wild TTS corpora generation. We apply a custom low-cost pipeline to the first in-the-wild Argentine Spanish collection and compare 24 pipeline configurations combining different denoising and quality filtering variants. Evaluation relies on complementary objective measures (PESQ, SI-SDR, SNR), acoustic descriptors (T30, C50), and speech-preservation metrics (F0-STD, MCD). Results expose trade-offs between dataset size, signal quality, and voice preservation; where denoising variants with permissive filtering provide the best overall compromise for our testbed. The proposed methodology allows selecting pipeline configurations without training TTS models for each subset, accelerating and reducing the cost of preprocessing development for low-resource settings. |
| title | Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2510.03111 |