A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yuewei, Zou, Huanbin, Zhu, Jie
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914646054666240
author Zhang, Yuewei
Zou, Huanbin
Zhu, Jie
author_facet Zhang, Yuewei
Zou, Huanbin
Zhu, Jie
contents Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and further modify the complex spectrum to suppress the residual noise and recover the speech phase in the second stage. The above whole process is performed in the short-time Fourier transform (STFT) spectrum domain. In this paper, we re-implement the above second sub-process in the short-time discrete cosine transform (STDCT) spectrum domain. The reason is that we have found STDCT performs greater noise suppression capability than STFT. Additionally, the implicit phase of STDCT ensures simpler and more efficient phase recovery, which is challenging and computationally expensive in the STFT-based methods. Therefore, we propose a novel two-stage framework called the STFT-STDCT spectrum fusion network (FDFNet) for speech enhancement in cross-spectrum domain. Experimental results demonstrate that the proposed FDFNet outperforms the previous two-stage methods and also exhibits superior performance compared to other advanced systems.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10494
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
Zhang, Yuewei
Zou, Huanbin
Zhu, Jie
Audio and Speech Processing
Sound
Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and further modify the complex spectrum to suppress the residual noise and recover the speech phase in the second stage. The above whole process is performed in the short-time Fourier transform (STFT) spectrum domain. In this paper, we re-implement the above second sub-process in the short-time discrete cosine transform (STDCT) spectrum domain. The reason is that we have found STDCT performs greater noise suppression capability than STFT. Additionally, the implicit phase of STDCT ensures simpler and more efficient phase recovery, which is challenging and computationally expensive in the STFT-based methods. Therefore, we propose a novel two-stage framework called the STFT-STDCT spectrum fusion network (FDFNet) for speech enhancement in cross-spectrum domain. Experimental results demonstrate that the proposed FDFNet outperforms the previous two-stage methods and also exhibits superior performance compared to other advanced systems.
title A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2401.10494