Autoregressive Speech Synthesis without Vector Quantization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Meng, Lingwei, Zhou, Long, Liu, Shujie, Chen, Sanyuan, Han, Bing, Hu, Shujie, Liu, Yanqing, Li, Jinyu, Zhao, Sheng, Wu, Xixin, Meng, Helen, Wei, Furu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916760052039680
author Meng, Lingwei
Zhou, Long
Liu, Shujie
Chen, Sanyuan
Han, Bing
Hu, Shujie
Liu, Yanqing
Li, Jinyu
Zhao, Sheng
Wu, Xixin
Meng, Helen
Wei, Furu
author_facet Meng, Lingwei
Zhou, Long
Liu, Shujie
Chen, Sanyuan
Han, Bing
Hu, Shujie
Liu, Yanqing
Li, Jinyu
Zhao, Sheng
Wu, Xixin
Meng, Helen
Wei, Furu
contents We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need for vector quantization, which is typically designed for audio compression and sacrifices fidelity compared to continuous representations. Specifically, (i) instead of cross-entropy loss, we apply regression loss with a proposed spectrogram flux loss function to model the probability distribution of the continuous-valued tokens; (ii) we have incorporated variational inference into MELLE to facilitate sampling mechanisms, thereby enhancing the output diversity and model robustness. Experiments demonstrate that, compared to the two-stage codec language model VALL-E and its variants, the single-stage MELLE mitigates robustness issues by avoiding the inherent flaws of sampling vector-quantized codes, achieves superior performance across multiple metrics, and, most importantly, offers a more streamlined paradigm. The demos of our work are provided at https://aka.ms/melle.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08551
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Autoregressive Speech Synthesis without Vector Quantization
Meng, Lingwei
Zhou, Long
Liu, Shujie
Chen, Sanyuan
Han, Bing
Hu, Shujie
Liu, Yanqing
Li, Jinyu
Zhao, Sheng
Wu, Xixin
Meng, Helen
Wei, Furu
Computation and Language
Sound
Audio and Speech Processing
We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need for vector quantization, which is typically designed for audio compression and sacrifices fidelity compared to continuous representations. Specifically, (i) instead of cross-entropy loss, we apply regression loss with a proposed spectrogram flux loss function to model the probability distribution of the continuous-valued tokens; (ii) we have incorporated variational inference into MELLE to facilitate sampling mechanisms, thereby enhancing the output diversity and model robustness. Experiments demonstrate that, compared to the two-stage codec language model VALL-E and its variants, the single-stage MELLE mitigates robustness issues by avoiding the inherent flaws of sampling vector-quantized codes, achieves superior performance across multiple metrics, and, most importantly, offers a more streamlined paradigm. The demos of our work are provided at https://aka.ms/melle.
title Autoregressive Speech Synthesis without Vector Quantization
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.08551