Towards Robust FastSpeech 2 by Modelling Residual Multimodality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kögel, Fabian, Nguyen, Bac, Cardinaux, Fabien
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914951157776384
author Kögel, Fabian
Nguyen, Bac
Cardinaux, Fabien
author_facet Kögel, Fabian
Nguyen, Bac
Cardinaux, Fabien
contents State-of-the-art non-autoregressive text-to-speech (TTS) models based on FastSpeech 2 can efficiently synthesise high-fidelity and natural speech. For expressive speech datasets however, we observe characteristic audio distortions. We demonstrate that such artefacts are introduced to the vocoder reconstruction by over-smooth mel-spectrogram predictions, which are induced by the choice of mean-squared-error (MSE) loss for training the mel-spectrogram decoder. With MSE loss FastSpeech 2 is limited to learn conditional averages of the training distribution, which might not lie close to a natural sample if the distribution still appears multimodal after all conditioning signals. To alleviate this problem, we introduce TVC-GMM, a mixture model of Trivariate-Chain Gaussian distributions, to model the residual multimodality. TVC-GMM reduces spectrogram smoothness and improves perceptual audio quality in particular for expressive datasets as shown by both objective and subjective evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2306_01442
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards Robust FastSpeech 2 by Modelling Residual Multimodality
Kögel, Fabian
Nguyen, Bac
Cardinaux, Fabien
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
State-of-the-art non-autoregressive text-to-speech (TTS) models based on FastSpeech 2 can efficiently synthesise high-fidelity and natural speech. For expressive speech datasets however, we observe characteristic audio distortions. We demonstrate that such artefacts are introduced to the vocoder reconstruction by over-smooth mel-spectrogram predictions, which are induced by the choice of mean-squared-error (MSE) loss for training the mel-spectrogram decoder. With MSE loss FastSpeech 2 is limited to learn conditional averages of the training distribution, which might not lie close to a natural sample if the distribution still appears multimodal after all conditioning signals. To alleviate this problem, we introduce TVC-GMM, a mixture model of Trivariate-Chain Gaussian distributions, to model the residual multimodality. TVC-GMM reduces spectrogram smoothness and improves perceptual audio quality in particular for expressive datasets as shown by both objective and subjective evaluation.
title Towards Robust FastSpeech 2 by Modelling Residual Multimodality
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2306.01442