SAM: A Mamba-2 State-Space Audio-Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Taehan, Jung, Jaehan, Lee, Hyukjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911485418012672
author Lee, Taehan
Jung, Jaehan
Lee, Hyukjun
author_facet Lee, Taehan
Jung, Jaehan
Lee, Hyukjun
contents We present SAM, a State-space Audio-language Model that integrates an audio encoder with a Mamba-2 backbone. SAM-2.7B achieves 21.1 mAP on AudioSet and 17.6 SPICE on AudioCaps, matching or surpassing larger 7B transformer-based models with fewer parameters. We further provide the first systematic, representation-level analysis of how SSMs interact with audio encoder outputs: (1) joint audio encoder finetuning is essential, supported by accuracy gains and observed adaptation of token representation rank and similarity across different SSM sizes; (2) despite linear scaling, SSMs benefit more from compact, information-rich audio token representations than from excessively long token sequences; and (3) incorporating instruction-following supervision substantially improves reasoning ability, boosting MMAU-Sound accuracy from 22.8 to 56.8. Through comprehensive experiments and analysis, we establish practical design principles for SSMs as strong, scalable backbones for audio-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAM: A Mamba-2 State-Space Audio-Language Model
Lee, Taehan
Jung, Jaehan
Lee, Hyukjun
Sound
Audio and Speech Processing
We present SAM, a State-space Audio-language Model that integrates an audio encoder with a Mamba-2 backbone. SAM-2.7B achieves 21.1 mAP on AudioSet and 17.6 SPICE on AudioCaps, matching or surpassing larger 7B transformer-based models with fewer parameters. We further provide the first systematic, representation-level analysis of how SSMs interact with audio encoder outputs: (1) joint audio encoder finetuning is essential, supported by accuracy gains and observed adaptation of token representation rank and similarity across different SSM sizes; (2) despite linear scaling, SSMs benefit more from compact, information-rich audio token representations than from excessively long token sequences; and (3) incorporating instruction-following supervision substantially improves reasoning ability, boosting MMAU-Sound accuracy from 22.8 to 56.8. Through comprehensive experiments and analysis, we establish practical design principles for SSMs as strong, scalable backbones for audio-language models.
title SAM: A Mamba-2 State-Space Audio-Language Model
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15680