wav2pos: Sound Source Localization using Masked Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929629512597504 |
|---|---|
| author | Berg, Axel Gulin, Jens O'Connor, Mark Zhou, Chuteng Åström, Karl Oskarsson, Magnus |
| author_facet | Berg, Axel Gulin, Jens O'Connor, Mark Zhou, Chuteng Åström, Karl Oskarsson, Magnus |
| contents | We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_15771 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | wav2pos: Sound Source Localization using Masked Autoencoders Berg, Axel Gulin, Jens O'Connor, Mark Zhou, Chuteng Åström, Karl Oskarsson, Magnus Audio and Speech Processing Machine Learning Sound We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods. |
| title | wav2pos: Sound Source Localization using Masked Autoencoders |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2408.15771 |