Binaural Localization Model for Speech in Noise

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tokala, Vikas, Grinstein, Eric, Brooks, Rory, Brookes, Mike, Doclo, Simon, Jensen, Jesper, Naylor, Patrick A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909708423528448
author Tokala, Vikas
Grinstein, Eric
Brooks, Rory
Brookes, Mike
Doclo, Simon
Jensen, Jesper
Naylor, Patrick A.
author_facet Tokala, Vikas
Grinstein, Eric
Brooks, Rory
Brookes, Mike
Doclo, Simon
Jensen, Jesper
Naylor, Patrick A.
contents Binaural acoustic source localization is important to human listeners for spatial awareness, communication and safety. In this paper, an end-to-end binaural localization model for speech in noise is presented. A lightweight convolutional recurrent network that localizes sound in the frontal azimuthal plane for noisy reverberant binaural signals is introduced. The model incorporates additive internal ear noise to represent the frequency-dependent hearing threshold of a typical listener. The localization performance of the model is compared with the steered response power algorithm, and the use of the model as a measure of interaural cue preservation for binaural speech enhancement methods is studied. A listening test was performed to compare the performance of the model with human localization of speech in noisy conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Binaural Localization Model for Speech in Noise
Tokala, Vikas
Grinstein, Eric
Brooks, Rory
Brookes, Mike
Doclo, Simon
Jensen, Jesper
Naylor, Patrick A.
Audio and Speech Processing
Signal Processing
Binaural acoustic source localization is important to human listeners for spatial awareness, communication and safety. In this paper, an end-to-end binaural localization model for speech in noise is presented. A lightweight convolutional recurrent network that localizes sound in the frontal azimuthal plane for noisy reverberant binaural signals is introduced. The model incorporates additive internal ear noise to represent the frequency-dependent hearing threshold of a typical listener. The localization performance of the model is compared with the steered response power algorithm, and the use of the model as a measure of interaural cue preservation for binaural speech enhancement methods is studied. A listening test was performed to compare the performance of the model with human localization of speech in noisy conditions.
title Binaural Localization Model for Speech in Noise
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2507.20027