A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Guoqiang, Tan, Huaning, Li, Ruilai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911950303133696
author Hu, Guoqiang
Tan, Huaning
Li, Ruilai
author_facet Hu, Guoqiang
Tan, Huaning
Li, Ruilai
contents Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its Fourier transform process, the clarity of speech synthesised by Mel spectrogram is compromised in mutant signals. In order to obtain a more detailed Mel spectrogram, we propose a Mel spectrogram enhancement paradigm based on the continuous wavelet transform (CWT). This paradigm introduces an additional task: a more detailed wavelet spectrogram, which like the post-processing network takes as input the Mel spectrogram output by the decoder. We choose Tacotron2 and Fastspeech2 for experimental validation in order to test autoregressive (AR) and non-autoregressive (NAR) speech systems, respectively. The experimental results demonstrate that the speech synthesised using the model with the Mel spectrogram enhancement paradigm exhibits higher MOS, with an improvement of 0.14 and 0.09 compared to the baseline model, respectively. These findings provide some validation for the universality of the enhancement paradigm, as they demonstrate the success of the paradigm in different architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
Hu, Guoqiang
Tan, Huaning
Li, Ruilai
Sound
Artificial Intelligence
Audio and Speech Processing
Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its Fourier transform process, the clarity of speech synthesised by Mel spectrogram is compromised in mutant signals. In order to obtain a more detailed Mel spectrogram, we propose a Mel spectrogram enhancement paradigm based on the continuous wavelet transform (CWT). This paradigm introduces an additional task: a more detailed wavelet spectrogram, which like the post-processing network takes as input the Mel spectrogram output by the decoder. We choose Tacotron2 and Fastspeech2 for experimental validation in order to test autoregressive (AR) and non-autoregressive (NAR) speech systems, respectively. The experimental results demonstrate that the speech synthesised using the model with the Mel spectrogram enhancement paradigm exhibits higher MOS, with an improvement of 0.14 and 0.09 compared to the baseline model, respectively. These findings provide some validation for the universality of the enhancement paradigm, as they demonstrate the success of the paradigm in different architectures.
title A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2406.12164