A comparative study of generative models for child voice conversion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sudro, Protima Nomo, Ragni, Anton, Hain, Thomas
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915673076137984
author Sudro, Protima Nomo
Ragni, Anton
Hain, Thomas
author_facet Sudro, Protima Nomo
Ragni, Anton
Hain, Thomas
contents Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in producing children speech and in particular adult to child VC has not been investigated. For adult to child VC, four generative models are compared: diffusion model, flow based model, variational autoencoders, and generative adversarial network. Results show that although converted speech outputs produce by those models appear plausible, they exhibit insufficient similarity with the target speaker characteristics. We introduce an efficient frequency warping technique that can be applied to the output of models, and which shows significant reduction of the mismatch between adult and child. The output of all the models are evaluated using both objective and subjective measures. In particular we compare specific speaker pairing using a unique corpus collected for dubbing of children speech.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12129
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A comparative study of generative models for child voice conversion
Sudro, Protima Nomo
Ragni, Anton
Hain, Thomas
Sound
Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in producing children speech and in particular adult to child VC has not been investigated. For adult to child VC, four generative models are compared: diffusion model, flow based model, variational autoencoders, and generative adversarial network. Results show that although converted speech outputs produce by those models appear plausible, they exhibit insufficient similarity with the target speaker characteristics. We introduce an efficient frequency warping technique that can be applied to the output of models, and which shows significant reduction of the mismatch between adult and child. The output of all the models are evaluated using both objective and subjective measures. In particular we compare specific speaker pairing using a unique corpus collected for dubbing of children speech.
title A comparative study of generative models for child voice conversion
topic Sound
url https://arxiv.org/abs/2512.12129