Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Diaz-Guerra, David, Politis, Archontis, Miguel, Antonio, Beltran, Jose R., Virtanen, Tuomas
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914693279383552
author Diaz-Guerra, David
Politis, Archontis
Miguel, Antonio
Beltran, Jose R.
Virtanen, Tuomas
author_facet Diaz-Guerra, David
Politis, Archontis
Miguel, Antonio
Beltran, Jose R.
Virtanen, Tuomas
contents Many multi-source localization and tracking models based on neural networks use one or several recurrent layers at their final stages to track the movement of the sources. Conventional recurrent neural networks (RNNs), such as the long short-term memories (LSTMs) or the gated recurrent units (GRUs), take a vector as their input and use another vector to store their state. However, this approach results in the information from all the sources being contained in a single ordered vector, which is not optimal for permutation-invariant problems such as multi-source tracking. In this paper, we present a new recurrent architecture that uses unordered sets to represent both its input and its state and that is invariant to the permutations of the input set and equivariant to the permutations of the state set. Hence, the information of every sound source is represented in an individual embedding and the new estimates are assigned to the tracked trajectories regardless of their order.
format Preprint
id arxiv_https___arxiv_org_abs_2306_08510
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications
Diaz-Guerra, David
Politis, Archontis
Miguel, Antonio
Beltran, Jose R.
Virtanen, Tuomas
Audio and Speech Processing
Machine Learning
Sound
Signal Processing
Many multi-source localization and tracking models based on neural networks use one or several recurrent layers at their final stages to track the movement of the sources. Conventional recurrent neural networks (RNNs), such as the long short-term memories (LSTMs) or the gated recurrent units (GRUs), take a vector as their input and use another vector to store their state. However, this approach results in the information from all the sources being contained in a single ordered vector, which is not optimal for permutation-invariant problems such as multi-source tracking. In this paper, we present a new recurrent architecture that uses unordered sets to represent both its input and its state and that is invariant to the permutations of the input set and equivariant to the permutations of the state set. Hence, the information of every sound source is represented in an individual embedding and the new estimates are assigned to the tracked trajectories regardless of their order.
title Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications
topic Audio and Speech Processing
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2306.08510