Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Di Bernardo, Matías, Misley, Emmanuel, Correa, Ignacio, Iacovelli, Mateo García, Mellino, Simón, Barrios, Gala Lucía Gonzalez
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914072703795200
author Di Bernardo, Matías
Misley, Emmanuel
Correa, Ignacio
Iacovelli, Mateo García
Mellino, Simón
Barrios, Gala Lucía Gonzalez
author_facet Di Bernardo, Matías
Misley, Emmanuel
Correa, Ignacio
Iacovelli, Mateo García
Mellino, Simón
Barrios, Gala Lucía Gonzalez
contents This work introduces a reproducible, metric-driven methodology to evaluate preprocessing pipelines for in-the-wild TTS corpora generation. We apply a custom low-cost pipeline to the first in-the-wild Argentine Spanish collection and compare 24 pipeline configurations combining different denoising and quality filtering variants. Evaluation relies on complementary objective measures (PESQ, SI-SDR, SNR), acoustic descriptors (T30, C50), and speech-preservation metrics (F0-STD, MCD). Results expose trade-offs between dataset size, signal quality, and voice preservation; where denoising variants with permissive filtering provide the best overall compromise for our testbed. The proposed methodology allows selecting pipeline configurations without training TTS models for each subset, accelerating and reducing the cost of preprocessing development for low-resource settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03111
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets
Di Bernardo, Matías
Misley, Emmanuel
Correa, Ignacio
Iacovelli, Mateo García
Mellino, Simón
Barrios, Gala Lucía Gonzalez
Audio and Speech Processing
This work introduces a reproducible, metric-driven methodology to evaluate preprocessing pipelines for in-the-wild TTS corpora generation. We apply a custom low-cost pipeline to the first in-the-wild Argentine Spanish collection and compare 24 pipeline configurations combining different denoising and quality filtering variants. Evaluation relies on complementary objective measures (PESQ, SI-SDR, SNR), acoustic descriptors (T30, C50), and speech-preservation metrics (F0-STD, MCD). Results expose trade-offs between dataset size, signal quality, and voice preservation; where denoising variants with permissive filtering provide the best overall compromise for our testbed. The proposed methodology allows selecting pipeline configurations without training TTS models for each subset, accelerating and reducing the cost of preprocessing development for low-resource settings.
title Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets
topic Audio and Speech Processing
url https://arxiv.org/abs/2510.03111