LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Yiwei, Li, Zhihan, Du, Chenpeng, Wang, Hankun, Chen, Xie, Yu, Kai
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916748656115712
author Guo, Yiwei
Li, Zhihan
Du, Chenpeng
Wang, Hankun
Chen, Xie
Yu, Kai
author_facet Guo, Yiwei
Li, Zhihan
Du, Chenpeng
Wang, Hankun
Chen, Xie
Yu, Kai
contents Although discrete speech tokens have exhibited strong potential for language model-based speech generation, their high bitrates and redundant timbre information restrict the development of such models. In this work, we propose LSCodec, a discrete speech codec that has both low bitrate and speaker decoupling ability. LSCodec adopts a multi-stage unsupervised training framework with a speaker perturbation technique. A continuous information bottleneck is first established, followed by vector quantization that produces a discrete speaker-decoupled space. A discrete token vocoder finally refines acoustic details from LSCodec. By reconstruction evaluations, LSCodec demonstrates superior intelligibility and audio quality with only a single codebook and smaller vocabulary size than baselines. Voice conversion and speaker probing experiments prove the excellent speaker disentanglement of LSCodec, and ablation study verifies the effectiveness of the proposed training framework.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15764
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
Guo, Yiwei
Li, Zhihan
Du, Chenpeng
Wang, Hankun
Chen, Xie
Yu, Kai
Audio and Speech Processing
Artificial Intelligence
Sound
Although discrete speech tokens have exhibited strong potential for language model-based speech generation, their high bitrates and redundant timbre information restrict the development of such models. In this work, we propose LSCodec, a discrete speech codec that has both low bitrate and speaker decoupling ability. LSCodec adopts a multi-stage unsupervised training framework with a speaker perturbation technique. A continuous information bottleneck is first established, followed by vector quantization that produces a discrete speaker-decoupled space. A discrete token vocoder finally refines acoustic details from LSCodec. By reconstruction evaluations, LSCodec demonstrates superior intelligibility and audio quality with only a single codebook and smaller vocabulary size than baselines. Voice conversion and speaker probing experiments prove the excellent speaker disentanglement of LSCodec, and ablation study verifies the effectiveness of the proposed training framework.
title LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2410.15764