Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dai, Lingling, Li, Andong, Lei, Tong, Yu, Meng, Li, Xiaodong, Zheng, Chengshi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908555220615168
author Dai, Lingling
Li, Andong
Lei, Tong
Yu, Meng
Li, Xiaodong
Zheng, Chengshi
author_facet Dai, Lingling
Li, Andong
Lei, Tong
Yu, Meng
Li, Xiaodong
Zheng, Chengshi
contents Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this paper, we start from two representative T-F domain vocoders, namely Vocos and APNet2, which belong to the single-stream and dual-stream modes for magnitude and phase estimation, respectively. When evaluating their performance on a large-scale dataset, we accidentally observe severe performance collapse of APNet2. To stabilize its performance, in this paper, we introduce three simple yet effective strategies, each targeting the topological space, the source space, and the output space, respectively. Specifically, we modify the architectural topology for better information exchange in the topological space, introduce prior knowledge to facilitate the generation process in the source space, and optimize the backpropagation process for parameter updates with an improved output format in the output space. Experimental results demonstrate that our proposed method effectively facilitates the joint estimation of magnitude and phase in APNet2, thus bridging the performance disparities between the single-stream and dual-stream vocoders.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
Dai, Lingling
Li, Andong
Lei, Tong
Yu, Meng
Li, Xiaodong
Zheng, Chengshi
Audio and Speech Processing
Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this paper, we start from two representative T-F domain vocoders, namely Vocos and APNet2, which belong to the single-stream and dual-stream modes for magnitude and phase estimation, respectively. When evaluating their performance on a large-scale dataset, we accidentally observe severe performance collapse of APNet2. To stabilize its performance, in this paper, we introduce three simple yet effective strategies, each targeting the topological space, the source space, and the output space, respectively. Specifically, we modify the architectural topology for better information exchange in the topological space, introduce prior knowledge to facilitate the generation process in the source space, and optimize the backpropagation process for parameter updates with an improved output format in the output space. Experimental results demonstrate that our proposed method effectively facilitates the joint estimation of magnitude and phase in APNet2, thus bridging the performance disparities between the single-stream and dual-stream vocoders.
title Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
topic Audio and Speech Processing
url https://arxiv.org/abs/2509.18806