Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915985960730624 |
|---|---|
| author | Lee, Dongheon Pandey, Ashutosh Parekh, Sanjeel Wong, Daniel Donley, Jacob Xu, Buye Azcarreta, Juan |
| author_facet | Lee, Dongheon Pandey, Ashutosh Parekh, Sanjeel Wong, Daniel Donley, Jacob Xu, Buye Azcarreta, Juan |
| contents | While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_04749 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement Lee, Dongheon Pandey, Ashutosh Parekh, Sanjeel Wong, Daniel Donley, Jacob Xu, Buye Azcarreta, Juan Audio and Speech Processing While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available. |
| title | Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2605.04749 |