Investigating self-supervised features for expressive, multilingual voice conversion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Martín-Cortinas, Álvaro, Sáez-Trigueros, Daniel, Beringer, Grzegorz, Vallés-Pérez, Iván, Barra-Chicote, Roberto, Tura-Vecino, Biel, Gabryś, Adam, Bilinski, Piotr, Merritt, Thomas, Lorenzo-Trueba, Jaime
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909608664104960
author Martín-Cortinas, Álvaro
Sáez-Trigueros, Daniel
Beringer, Grzegorz
Vallés-Pérez, Iván
Barra-Chicote, Roberto
Tura-Vecino, Biel
Gabryś, Adam
Bilinski, Piotr
Merritt, Thomas
Lorenzo-Trueba, Jaime
author_facet Martín-Cortinas, Álvaro
Sáez-Trigueros, Daniel
Beringer, Grzegorz
Vallés-Pérez, Iván
Barra-Chicote, Roberto
Tura-Vecino, Biel
Gabryś, Adam
Bilinski, Piotr
Merritt, Thomas
Lorenzo-Trueba, Jaime
contents Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to produce. Unsupervised approaches are typically trained to reconstruct the input signal, which is composed of the content and the speaker information. Disentangling these components is a challenge and often leads to speaker leakage or prosodic information removal. In this paper, we explore voice conversion by leveraging the potential of self-supervised learning (SSL). A combination of the latent representations of SSL models, concatenated with speaker embeddings, is fed to a vocoder which is trained to reconstruct the input. Zero-shot voice conversion results show that this approach allows to keep the prosody and content of the source speaker while matching the speaker similarity of a VC system based on phonetic posteriorgrams (PPGs).
format Preprint
id arxiv_https___arxiv_org_abs_2505_08278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating self-supervised features for expressive, multilingual voice conversion
Martín-Cortinas, Álvaro
Sáez-Trigueros, Daniel
Beringer, Grzegorz
Vallés-Pérez, Iván
Barra-Chicote, Roberto
Tura-Vecino, Biel
Gabryś, Adam
Bilinski, Piotr
Merritt, Thomas
Lorenzo-Trueba, Jaime
Audio and Speech Processing
Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to produce. Unsupervised approaches are typically trained to reconstruct the input signal, which is composed of the content and the speaker information. Disentangling these components is a challenge and often leads to speaker leakage or prosodic information removal. In this paper, we explore voice conversion by leveraging the potential of self-supervised learning (SSL). A combination of the latent representations of SSL models, concatenated with speaker embeddings, is fed to a vocoder which is trained to reconstruct the input. Zero-shot voice conversion results show that this approach allows to keep the prosody and content of the source speaker while matching the speaker similarity of a VC system based on phonetic posteriorgrams (PPGs).
title Investigating self-supervised features for expressive, multilingual voice conversion
topic Audio and Speech Processing
url https://arxiv.org/abs/2505.08278