Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Du, Hui-Peng, Ai, Yang, Zheng, Rui-Chen, Lu, Ye-Xin, Ling, Zhen-Hua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908484328488960
author Du, Hui-Peng
Ai, Yang
Zheng, Rui-Chen
Lu, Ye-Xin
Ling, Zhen-Hua
author_facet Du, Hui-Peng
Ai, Yang
Zheng, Rui-Chen
Lu, Ye-Xin
Ling, Zhen-Hua
contents Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model complexity. Therefore, this paper proposes a novel FreeGAN vocoder, aiming to answer the question of whether GAN is necessary for mel-spectrogram-based neural vocoders. The FreeGAN employs an amplitude-phase serial prediction framework, eliminating the need for GAN training. It incorporates amplitude prior input, SNAKE-ConvNeXt v2 backbone and frequency-weighted anti-wrapping phase loss to compensate for the performance loss caused by the absence of GAN. Experimental results confirm that the speech quality of FreeGAN is comparable to that of advanced GAN-based vocoders, while significantly improving training efficiency and complexity. Other explicit-phase-prediction-based neural vocoders can also work without GAN, leveraging our proposed methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
Du, Hui-Peng
Ai, Yang
Zheng, Rui-Chen
Lu, Ye-Xin
Ling, Zhen-Hua
Audio and Speech Processing
Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model complexity. Therefore, this paper proposes a novel FreeGAN vocoder, aiming to answer the question of whether GAN is necessary for mel-spectrogram-based neural vocoders. The FreeGAN employs an amplitude-phase serial prediction framework, eliminating the need for GAN training. It incorporates amplitude prior input, SNAKE-ConvNeXt v2 backbone and frequency-weighted anti-wrapping phase loss to compensate for the performance loss caused by the absence of GAN. Experimental results confirm that the speech quality of FreeGAN is comparable to that of advanced GAN-based vocoders, while significantly improving training efficiency and complexity. Other explicit-phase-prediction-based neural vocoders can also work without GAN, leveraging our proposed methods.
title Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.07711