Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Dongheon, Pandey, Ashutosh, Parekh, Sanjeel, Wong, Daniel, Donley, Jacob, Xu, Buye, Azcarreta, Juan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915985960730624
author Lee, Dongheon
Pandey, Ashutosh
Parekh, Sanjeel
Wong, Daniel
Donley, Jacob
Xu, Buye
Azcarreta, Juan
author_facet Lee, Dongheon
Pandey, Ashutosh
Parekh, Sanjeel
Wong, Daniel
Donley, Jacob
Xu, Buye
Azcarreta, Juan
contents While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
Lee, Dongheon
Pandey, Ashutosh
Parekh, Sanjeel
Wong, Daniel
Donley, Jacob
Xu, Buye
Azcarreta, Juan
Audio and Speech Processing
While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available.
title Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
topic Audio and Speech Processing
url https://arxiv.org/abs/2605.04749