Unsupervised Speech Enhancement using Data-defined Priors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Klement, Dominik, Maciejewski, Matthew, Khudanpur, Sanjeev, Černocký, Jan, Burget, Lukáš
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916973229637632
author Klement, Dominik
Maciejewski, Matthew
Khudanpur, Sanjeev
Černocký, Jan
Burget, Lukáš
author_facet Klement, Dominik
Maciejewski, Matthew
Khudanpur, Sanjeev
Černocký, Jan
Burget, Lukáš
contents The majority of deep learning-based speech enhancement methods require paired clean-noisy speech data. Collecting such data at scale in real-world conditions is infeasible, which has led the community to rely on synthetically generated noisy speech. However, this introduces a gap between the training and testing phases. In this work, we propose a novel dual-branch encoder-decoder architecture for unsupervised speech enhancement that separates the input into clean speech and residual noise. Adversarial training is employed to impose priors on each branch, defined by unpaired datasets of clean speech and, optionally, noise. Experimental results show that our method achieves performance comparable to leading unsupervised speech enhancement approaches. Furthermore, we demonstrate the critical impact of clean speech data selection on enhancement performance. In particular, our findings reveal that performance may appear overly optimistic when in-domain clean speech data are used for prior definition -- a practice adopted in previous unsupervised speech enhancement studies.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Speech Enhancement using Data-defined Priors
Klement, Dominik
Maciejewski, Matthew
Khudanpur, Sanjeev
Černocký, Jan
Burget, Lukáš
Audio and Speech Processing
Artificial Intelligence
Sound
The majority of deep learning-based speech enhancement methods require paired clean-noisy speech data. Collecting such data at scale in real-world conditions is infeasible, which has led the community to rely on synthetically generated noisy speech. However, this introduces a gap between the training and testing phases. In this work, we propose a novel dual-branch encoder-decoder architecture for unsupervised speech enhancement that separates the input into clean speech and residual noise. Adversarial training is employed to impose priors on each branch, defined by unpaired datasets of clean speech and, optionally, noise. Experimental results show that our method achieves performance comparable to leading unsupervised speech enhancement approaches. Furthermore, we demonstrate the critical impact of clean speech data selection on enhancement performance. In particular, our findings reveal that performance may appear overly optimistic when in-domain clean speech data are used for prior definition -- a practice adopted in previous unsupervised speech enhancement studies.
title Unsupervised Speech Enhancement using Data-defined Priors
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2509.22942