Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tokala, Vikas, Grinstein, Eric, Brookes, Mike, Doclo, Simon, Jensen, Jesper, Naylor, Patrick A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913961892380672
author Tokala, Vikas
Grinstein, Eric
Brookes, Mike
Doclo, Simon
Jensen, Jesper
Naylor, Patrick A.
author_facet Tokala, Vikas
Grinstein, Eric
Brookes, Mike
Doclo, Simon
Jensen, Jesper
Naylor, Patrick A.
contents From hearing aids to augmented and virtual reality devices, binaural speech enhancement algorithms have been established as state-of-the-art techniques to improve speech intelligibility and listening comfort. In this paper, we present an end-to-end binaural speech enhancement method using a complex recurrent convolutional network with an encoder-decoder architecture and a complex LSTM recurrent block placed between the encoder and decoder. A loss function that focuses on the preservation of spatial information in addition to speech intelligibility improvement and noise reduction is introduced. The network estimates individual complex ratio masks for the left and right-ear channels of a binaural hearing device in the time-frequency domain. We show that, compared to other baseline algorithms, the proposed method significantly improves the estimated speech intelligibility and reduces the noise while preserving the spatial information of the binaural signals in acoustic situations with a single target speaker and isotropic noise of various types.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20023
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
Tokala, Vikas
Grinstein, Eric
Brookes, Mike
Doclo, Simon
Jensen, Jesper
Naylor, Patrick A.
Audio and Speech Processing
Signal Processing
From hearing aids to augmented and virtual reality devices, binaural speech enhancement algorithms have been established as state-of-the-art techniques to improve speech intelligibility and listening comfort. In this paper, we present an end-to-end binaural speech enhancement method using a complex recurrent convolutional network with an encoder-decoder architecture and a complex LSTM recurrent block placed between the encoder and decoder. A loss function that focuses on the preservation of spatial information in addition to speech intelligibility improvement and noise reduction is introduced. The network estimates individual complex ratio masks for the left and right-ear channels of a binaural hearing device in the time-frequency domain. We show that, compared to other baseline algorithms, the proposed method significantly improves the estimated speech intelligibility and reduces the noise while preserving the spatial information of the binaural signals in acoustic situations with a single target speaker and isotropic noise of various types.
title Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2507.20023