State-Space Large Audio Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bhati, Saurabhchand, Gong, Yuan, Karlinsky, Leonid, Kuehne, Hilde, Feris, Rogerio, Glass, James
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915032250449920
author Bhati, Saurabhchand
Gong, Yuan
Karlinsky, Leonid
Kuehne, Hilde
Feris, Rogerio
Glass, James
author_facet Bhati, Saurabhchand
Gong, Yuan
Karlinsky, Leonid
Kuehne, Hilde
Feris, Rogerio
Glass, James
contents Large Audio Language Models (LALM) combine the audio perception models and the Large Language Models (LLM) and show a remarkable ability to reason about the input audio, infer the meaning, and understand the intent. However, these systems rely on Transformers which scale quadratically with the input sequence lengths which poses computational challenges in deploying these systems in memory and time-constrained scenarios. Recently, the state-space models (SSMs) have emerged as an alternative to transformer networks. While there have been successful attempts to replace transformer-based audio perception models with state-space ones, state-space-based LALMs remain unexplored. First, we begin by replacing the transformer-based audio perception module and then replace the transformer-based LLM and propose the first state-space-based LALM. Experimental results demonstrate that space-based LALM despite having a significantly lower number of parameters performs competitively with transformer-based LALMs on close-ended tasks on a variety of datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15685
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle State-Space Large Audio Language Models
Bhati, Saurabhchand
Gong, Yuan
Karlinsky, Leonid
Kuehne, Hilde
Feris, Rogerio
Glass, James
Audio and Speech Processing
Artificial Intelligence
Large Audio Language Models (LALM) combine the audio perception models and the Large Language Models (LLM) and show a remarkable ability to reason about the input audio, infer the meaning, and understand the intent. However, these systems rely on Transformers which scale quadratically with the input sequence lengths which poses computational challenges in deploying these systems in memory and time-constrained scenarios. Recently, the state-space models (SSMs) have emerged as an alternative to transformer networks. While there have been successful attempts to replace transformer-based audio perception models with state-space ones, state-space-based LALMs remain unexplored. First, we begin by replacing the transformer-based audio perception module and then replace the transformer-based LLM and propose the first state-space-based LALM. Experimental results demonstrate that space-based LALM despite having a significantly lower number of parameters performs competitively with transformer-based LALMs on close-ended tasks on a variety of datasets.
title State-Space Large Audio Language Models
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2411.15685