Explainable DNN-based Beamformer with Postfilter

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cohen, Adi, Wong, Daniel, Lee, Jung-Suk, Gannot, Sharon
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915024243523584
author Cohen, Adi
Wong, Daniel
Lee, Jung-Suk
Gannot, Sharon
author_facet Cohen, Adi
Wong, Daniel
Lee, Jung-Suk
Gannot, Sharon
contents This paper introduces an explainable DNN-based beamformer with a postfilter (ExNet-BF+PF) for multichannel signal processing. Our approach combines the U-Net network with a beamformer structure to address this problem. The method involves a two-stage processing pipeline. In the first stage, time-invariant weights are applied to construct a multichannel spatial filter, namely a beamformer. In the second stage, a time-varying single-channel post-filter is applied at the beamformer output. Additionally, we incorporate an attention mechanism inspired by its successful application in noisy and reverberant environments to improve speech enhancement further. Furthermore, our study fills a gap in the existing literature by conducting a thorough spatial analysis of the network's performance. Specifically, we examine how the network utilizes spatial information during processing. This analysis yields valuable insights into the network's functionality, thereby enhancing our understanding of its overall performance. Experimental results demonstrate that our approach is not only straightforward to train but also yields superior results, obviating the necessity for prior knowledge of the speaker's activity.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10854
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Explainable DNN-based Beamformer with Postfilter
Cohen, Adi
Wong, Daniel
Lee, Jung-Suk
Gannot, Sharon
Audio and Speech Processing
This paper introduces an explainable DNN-based beamformer with a postfilter (ExNet-BF+PF) for multichannel signal processing. Our approach combines the U-Net network with a beamformer structure to address this problem. The method involves a two-stage processing pipeline. In the first stage, time-invariant weights are applied to construct a multichannel spatial filter, namely a beamformer. In the second stage, a time-varying single-channel post-filter is applied at the beamformer output. Additionally, we incorporate an attention mechanism inspired by its successful application in noisy and reverberant environments to improve speech enhancement further. Furthermore, our study fills a gap in the existing literature by conducting a thorough spatial analysis of the network's performance. Specifically, we examine how the network utilizes spatial information during processing. This analysis yields valuable insights into the network's functionality, thereby enhancing our understanding of its overall performance. Experimental results demonstrate that our approach is not only straightforward to train but also yields superior results, obviating the necessity for prior knowledge of the speaker's activity.
title Explainable DNN-based Beamformer with Postfilter
topic Audio and Speech Processing
url https://arxiv.org/abs/2411.10854