Audio Mamba: Bidirectional State Space Model for Audio Representation Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Erol, Mehmet Hamza, Senocak, Arda, Feng, Jiu, Chung, Joon Son
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914825345433600
author Erol, Mehmet Hamza
Senocak, Arda
Feng, Jiu
Chung, Joon Son
author_facet Erol, Mehmet Hamza
Senocak, Arda
Feng, Jiu
Chung, Joon Son
contents Transformers have rapidly become the preferred choice for audio classification, surpassing methods based on CNNs. However, Audio Spectrogram Transformers (ASTs) exhibit quadratic scaling due to self-attention. The removal of this quadratic self-attention cost presents an appealing direction. Recently, state space models (SSMs), such as Mamba, have demonstrated potential in language and vision tasks in this regard. In this study, we explore whether reliance on self-attention is necessary for audio classification tasks. By introducing Audio Mamba (AuM), the first self-attention-free, purely SSM-based model for audio classification, we aim to address this question. We evaluate AuM on various audio datasets - comprising six different benchmarks - where it achieves comparable or better performance compared to well-established AST model.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03344
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Erol, Mehmet Hamza
Senocak, Arda
Feng, Jiu
Chung, Joon Son
Sound
Artificial Intelligence
Audio and Speech Processing
Transformers have rapidly become the preferred choice for audio classification, surpassing methods based on CNNs. However, Audio Spectrogram Transformers (ASTs) exhibit quadratic scaling due to self-attention. The removal of this quadratic self-attention cost presents an appealing direction. Recently, state space models (SSMs), such as Mamba, have demonstrated potential in language and vision tasks in this regard. In this study, we explore whether reliance on self-attention is necessary for audio classification tasks. By introducing Audio Mamba (AuM), the first self-attention-free, purely SSM-based model for audio classification, we aim to address this question. We evaluate AuM on various audio datasets - comprising six different benchmarks - where it achieves comparable or better performance compared to well-established AST model.
title Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2406.03344