Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Akti, Seymanur, Nguyen, Tuan Nam, Waibel, Alexander
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910987047665664
author Akti, Seymanur
Nguyen, Tuan Nam
Waibel, Alexander
author_facet Akti, Seymanur
Nguyen, Tuan Nam
Waibel, Alexander
contents Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional variational autoencoder, focusing on reducing source timbre leakage and improving linguistic-acoustic disentanglement for better style transfer. To minimize style leakage, we use multilingual discrete speech units for content representation and reinforce embeddings with augmentation-based similarity loss and mix-style layer normalization. To enhance expressivity transfer, we incorporate local F0 information via cross-attention and extract style embeddings enriched with global pitch and energy features. Experiments show our model outperforms baselines in emotion and speaker similarity, demonstrating superior style adaptation and reduced source style leakage.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04013
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
Akti, Seymanur
Nguyen, Tuan Nam
Waibel, Alexander
Sound
Artificial Intelligence
Audio and Speech Processing
Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional variational autoencoder, focusing on reducing source timbre leakage and improving linguistic-acoustic disentanglement for better style transfer. To minimize style leakage, we use multilingual discrete speech units for content representation and reinforce embeddings with augmentation-based similarity loss and mix-style layer normalization. To enhance expressivity transfer, we incorporate local F0 information via cross-attention and extract style embeddings enriched with global pitch and energy features. Experiments show our model outperforms baselines in emotion and speaker similarity, demonstrating superior style adaptation and reduced source style leakage.
title Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.04013