wav2pos: Sound Source Localization using Masked Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berg, Axel, Gulin, Jens, O'Connor, Mark, Zhou, Chuteng, Åström, Karl, Oskarsson, Magnus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929629512597504
author Berg, Axel
Gulin, Jens
O'Connor, Mark
Zhou, Chuteng
Åström, Karl
Oskarsson, Magnus
author_facet Berg, Axel
Gulin, Jens
O'Connor, Mark
Zhou, Chuteng
Åström, Karl
Oskarsson, Magnus
contents We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15771
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle wav2pos: Sound Source Localization using Masked Autoencoders
Berg, Axel
Gulin, Jens
O'Connor, Mark
Zhou, Chuteng
Åström, Karl
Oskarsson, Magnus
Audio and Speech Processing
Machine Learning
Sound
We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio recordings and microphone coordinates, we show that such a formulation allows for accurate localization of the sound source, by reconstructing coordinates masked in the input. Our approach is flexible in the sense that a single model can be used with an arbitrary number of microphones, even when a subset of audio recordings and microphone coordinates are missing. We test our method on simulated and real-world recordings of music and speech in indoor environments, and demonstrate competitive performance compared to both classical and other learning based localization methods.
title wav2pos: Sound Source Localization using Masked Autoencoders
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2408.15771