Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mittal, Manan, Deppisch, Thomas, Forrer, Joseph, Sueur, Chris Le, Ben-Hur, Zamir, Alon, David Lou, Wong, Daniel D. E.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911674674446336
author Mittal, Manan
Deppisch, Thomas
Forrer, Joseph
Sueur, Chris Le
Ben-Hur, Zamir
Alon, David Lou
Wong, Daniel D. E.
author_facet Mittal, Manan
Deppisch, Thomas
Forrer, Joseph
Sueur, Chris Le
Ben-Hur, Zamir
Alon, David Lou
Wong, Daniel D. E.
contents We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13548
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers
Mittal, Manan
Deppisch, Thomas
Forrer, Joseph
Sueur, Chris Le
Ben-Hur, Zamir
Alon, David Lou
Wong, Daniel D. E.
Sound
Audio and Speech Processing
Machine Learning
We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress sounds from selected directions while preserving natural binaural cues. Unlike traditional methods that rely on explicit direction-of-arrival estimation or operate in the Ambisonics domain, our signal-dependent framework combines multiple binaural filters in an online manner using implicit localization. This allows for real-time tracking and enhancement of moving sound sources, supporting applications such as speech focus, noise reduction, and world-locked audio in augmented and virtual reality. The method is agnostic to array geometry offering a flexible solution for spatial audio capture and personalized playback in next-generation consumer audio devices.
title Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers
topic Sound
Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2509.13548