Continuous Speech Tokenizer in Text To Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yixing, Xie, Ruobing, Sun, Xingwu, Cheng, Yu, Kang, Zhanhui
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917972282441728
author Li, Yixing
Xie, Ruobing
Sun, Xingwu
Cheng, Yu
Kang, Zhanhui
author_facet Li, Yixing
Xie, Ruobing
Sun, Xingwu
Cheng, Yu
Kang, Zhanhui
contents The fusion of speech and language in the era of large language models has garnered significant attention. Discrete speech token is often utilized in text-to-speech tasks for speech compression and portability, which is convenient for joint training with text and have good compression efficiency. However, we found that the discrete speech tokenizer still suffers from information loss. Therefore, we propose a simple yet effective continuous speech tokenizer named Cont-SPT, and a text-to-speech model based on continuous speech tokens. Our results show that the speech language model based on the continuous speech tokenizer has better continuity and higher estimated Mean Opinion Scores (MoS). This enhancement is attributed to better information preservation rate of the continuous speech tokenizer across both low and high frequencies in the frequency domain. The code and resources for Cont-SPT can be found in https://github.com/Yixing-Li/Continuous-Speech-Tokenizer
format Preprint
id arxiv_https___arxiv_org_abs_2410_17081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous Speech Tokenizer in Text To Speech
Li, Yixing
Xie, Ruobing
Sun, Xingwu
Cheng, Yu
Kang, Zhanhui
Sound
Computation and Language
Audio and Speech Processing
The fusion of speech and language in the era of large language models has garnered significant attention. Discrete speech token is often utilized in text-to-speech tasks for speech compression and portability, which is convenient for joint training with text and have good compression efficiency. However, we found that the discrete speech tokenizer still suffers from information loss. Therefore, we propose a simple yet effective continuous speech tokenizer named Cont-SPT, and a text-to-speech model based on continuous speech tokens. Our results show that the speech language model based on the continuous speech tokenizer has better continuity and higher estimated Mean Opinion Scores (MoS). This enhancement is attributed to better information preservation rate of the continuous speech tokenizer across both low and high frequencies in the frequency domain. The code and resources for Cont-SPT can be found in https://github.com/Yixing-Li/Continuous-Speech-Tokenizer
title Continuous Speech Tokenizer in Text To Speech
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2410.17081