Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruggiero, Giuseppe, Testa, Matteo, Van de Walle, Jurgen, Di Caro, Luigi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914957652656128
author Ruggiero, Giuseppe
Testa, Matteo
Van de Walle, Jurgen
Di Caro, Luigi
author_facet Ruggiero, Giuseppe
Testa, Matteo
Van de Walle, Jurgen
Di Caro, Luigi
contents The creation of artificial polyglot voices remains a challenging task, despite considerable progress in recent years. This paper investigates self-supervised learning for voice conversion to create native-sounding polyglot voices. We introduce a novel cross-lingual any-to-one voice conversion system that is able to preserve the source accent without the need for multilingual data from the target speaker. In addition, we show a novel cross-lingual fine-tuning strategy that further improves the accent and reduces the training data requirements. Objective and subjective evaluations with English, Spanish, French and Mandarin Chinese confirm that our approach improves on state-of-the-art methods, enhancing the speech intelligibility and overall quality of the converted speech, especially in cross-lingual scenarios. Audio samples are available at https://giuseppe-ruggiero.github.io/a2o-vc-demo/
format Preprint
id arxiv_https___arxiv_org_abs_2409_17387
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
Ruggiero, Giuseppe
Testa, Matteo
Van de Walle, Jurgen
Di Caro, Luigi
Sound
Audio and Speech Processing
The creation of artificial polyglot voices remains a challenging task, despite considerable progress in recent years. This paper investigates self-supervised learning for voice conversion to create native-sounding polyglot voices. We introduce a novel cross-lingual any-to-one voice conversion system that is able to preserve the source accent without the need for multilingual data from the target speaker. In addition, we show a novel cross-lingual fine-tuning strategy that further improves the accent and reduces the training data requirements. Objective and subjective evaluations with English, Spanish, French and Mandarin Chinese confirm that our approach improves on state-of-the-art methods, enhancing the speech intelligibility and overall quality of the converted speech, especially in cross-lingual scenarios. Audio samples are available at https://giuseppe-ruggiero.github.io/a2o-vc-demo/
title Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.17387