GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shetu, Shrishti Saha, Habets, Emanuël A. P., Brendel, Andreas
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912075444387840
author Shetu, Shrishti Saha
Habets, Emanuël A. P.
Brendel, Andreas
author_facet Shetu, Shrishti Saha
Habets, Emanuël A. P.
Brendel, Andreas
contents Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-arts discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality
format Preprint
id arxiv_https___arxiv_org_abs_2410_13599
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
Shetu, Shrishti Saha
Habets, Emanuël A. P.
Brendel, Andreas
Audio and Speech Processing
Sound
Signal Processing
Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-arts discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality
title GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2410.13599