JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cho, Hyunjae, Lee, Junhyeok, Jung, Wonbin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909220276797440
author Cho, Hyunjae
Lee, Junhyeok
Jung, Wonbin
author_facet Cho, Hyunjae
Lee, Junhyeok
Jung, Wonbin
contents Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we propose JenGAN, a new training strategy that involves stacking shifted low-pass filters to ensure the shift-equivariant property. This method helps prevent aliasing and reduce artifacts while preserving the model structure used during inference. In our experimental evaluation, JenGAN consistently enhances the performance of vocoder models, yielding significantly superior scores across the majority of evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06111
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
Cho, Hyunjae
Lee, Junhyeok
Jung, Wonbin
Audio and Speech Processing
Artificial Intelligence
Sound
Signal Processing
Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we propose JenGAN, a new training strategy that involves stacking shifted low-pass filters to ensure the shift-equivariant property. This method helps prevent aliasing and reduce artifacts while preserving the model structure used during inference. In our experimental evaluation, JenGAN consistently enhances the performance of vocoder models, yielding significantly superior scores across the majority of evaluation metrics.
title JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
topic Audio and Speech Processing
Artificial Intelligence
Sound
Signal Processing
url https://arxiv.org/abs/2406.06111