Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Yimin, Tang, Huaizhen, Zhang, Xulong, Cheng, Ning, Xiao, Jing, Wang, Jianzong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929213968220160
author Deng, Yimin
Tang, Huaizhen
Zhang, Xulong
Cheng, Ning
Xiao, Jing
Wang, Jianzong
author_facet Deng, Yimin
Tang, Huaizhen
Zhang, Xulong
Cheng, Ning
Xiao, Jing
Wang, Jianzong
contents Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides, the speaker-style modeling with pre-trained models making the process more complex. To tackle these issues, we introduce a new method named "CTVC" which utilizes disentangled speech representations with contrastive learning and time-invariant retrieval. Specifically, a similarity-based compression module is used to facilitate a more intimate connection between the frame-level hidden features and linguistic information at phoneme-level. Additionally, a time-invariant retrieval is proposed for timbre extraction based on multiple segmentations and mutual information. Experimental results demonstrate that "CTVC" outperforms previous studies and improves the sound quality and similarity of converted results.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08096
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
Deng, Yimin
Tang, Huaizhen
Zhang, Xulong
Cheng, Ning
Xiao, Jing
Wang, Jianzong
Sound
Audio and Speech Processing
Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides, the speaker-style modeling with pre-trained models making the process more complex. To tackle these issues, we introduce a new method named "CTVC" which utilizes disentangled speech representations with contrastive learning and time-invariant retrieval. Specifically, a similarity-based compression module is used to facilitate a more intimate connection between the frame-level hidden features and linguistic information at phoneme-level. Additionally, a time-invariant retrieval is proposed for timbre extraction based on multiple segmentations and mutual information. Experimental results demonstrate that "CTVC" outperforms previous studies and improves the sound quality and similarity of converted results.
title Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.08096