Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fiorio, Luan Vinícius, Defraene, Bruno, David, Johan, Young, Alex, Widdershoven, Frans, van Houtum, Wim, Aarts, Ronald M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916660749795328
author Fiorio, Luan Vinícius
Defraene, Bruno
David, Johan
Young, Alex
Widdershoven, Frans
van Houtum, Wim
Aarts, Ronald M.
author_facet Fiorio, Luan Vinícius
Defraene, Bruno
David, Johan
Young, Alex
Widdershoven, Frans
van Houtum, Wim
Aarts, Ronald M.
contents We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows the neural network model to learn toward which speaker to focus, during training, in a multi-speaker scenario, based on the position of listener and speakers. However, only audio information is necessary during inference. We perform acoustic simulations demonstrating the feasibility and performance when the SSM is employed in training. The results show significant increase in speech intelligibility, quality, and distortion metrics when compared to the minimum variance distortionless filter and the same neural network model trained without SSM. The success of the proposed method is a significant step forward toward the solution of the cocktail party problem.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios
Fiorio, Luan Vinícius
Defraene, Bruno
David, Johan
Young, Alex
Widdershoven, Frans
van Houtum, Wim
Aarts, Ronald M.
Audio and Speech Processing
Signal Processing
We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows the neural network model to learn toward which speaker to focus, during training, in a multi-speaker scenario, based on the position of listener and speakers. However, only audio information is necessary during inference. We perform acoustic simulations demonstrating the feasibility and performance when the SSM is employed in training. The results show significant increase in speech intelligibility, quality, and distortion metrics when compared to the minimum variance distortionless filter and the same neural network model trained without SSM. The success of the proposed method is a significant step forward toward the solution of the cocktail party problem.
title Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2503.18590