Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Yimin, Wang, Jianzong, Zhang, Xulong, Cheng, Ning, Xiao, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913338285359104
author Deng, Yimin
Wang, Jianzong
Zhang, Xulong
Cheng, Ning
Xiao, Jing
author_facet Deng, Yimin
Wang, Jianzong
Zhang, Xulong
Cheng, Ning
Xiao, Jing
contents Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in these representations, a lot of hidden speaker information leads to timbre leakage while the prosodic information of hidden units lacks use. To address these issues, we propose a novel framework for expressive voice conversion called "SAVC" based on soft speech units from HuBert-soft. Taking soft speech units as input, we design an attribute encoder to extract content and prosody features respectively. Specifically, we first introduce statistic perturbation imposed by adversarial style augmentation to eliminate speaker information. Then the prosody is implicitly modeled on soft speech units with knowledge distillation. Experiment results show that the intelligibility and naturalness of converted speech outperform previous work.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00603
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
Deng, Yimin
Wang, Jianzong
Zhang, Xulong
Cheng, Ning
Xiao, Jing
Sound
Audio and Speech Processing
Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in these representations, a lot of hidden speaker information leads to timbre leakage while the prosodic information of hidden units lacks use. To address these issues, we propose a novel framework for expressive voice conversion called "SAVC" based on soft speech units from HuBert-soft. Taking soft speech units as input, we design an attribute encoder to extract content and prosody features respectively. Specifically, we first introduce statistic perturbation imposed by adversarial style augmentation to eliminate speaker information. Then the prosody is implicitly modeled on soft speech units with knowledge distillation. Experiment results show that the intelligibility and naturalness of converted speech outperform previous work.
title Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.00603