Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yadav, Sarthak, Tan, Zheng-Hua
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929378263302144
author Yadav, Sarthak
Tan, Zheng-Hua
author_facet Yadav, Sarthak
Tan, Zheng-Hua
contents Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which have demonstrated promising results for language modelling. However, their feasibility for learning self-supervised, general-purpose audio representations is yet to be investigated. This work proposes Audio Mamba, a selective state space model for learning general-purpose audio representations from randomly masked spectrogram patches through self-supervision. Empirical results on ten diverse audio recognition downstream tasks show that the proposed models, pretrained on the AudioSet dataset, consistently outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by a considerable margin and demonstrate better performance in dataset size, sequence length and model size comparisons.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
Yadav, Sarthak
Tan, Zheng-Hua
Sound
Artificial Intelligence
Audio and Speech Processing
Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which have demonstrated promising results for language modelling. However, their feasibility for learning self-supervised, general-purpose audio representations is yet to be investigated. This work proposes Audio Mamba, a selective state space model for learning general-purpose audio representations from randomly masked spectrogram patches through self-supervision. Empirical results on ten diverse audio recognition downstream tasks show that the proposed models, pretrained on the AudioSet dataset, consistently outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by a considerable margin and demonstrate better performance in dataset size, sequence length and model size comparisons.
title Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2406.02178