Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Masuyama, Yoshiki, Miyazaki, Koichi, Murata, Masato
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909383702609920
author Masuyama, Yoshiki
Miyazaki, Koichi
Murata, Masato
author_facet Masuyama, Yoshiki
Miyazaki, Koichi
Murata, Masato
contents Selective state space models (SSMs) represented by Mamba have demonstrated their computational efficiency and promising outcomes in various tasks, including automatic speech recognition (ASR). Mamba has been applied to ASR task with the attention-based encoder-decoder framework, where the cross-attention mechanism between encoder and decoder remains. This paper explores the capability of Mamba as the decoder-only architecture in ASR task. Our MAmba-based DEcoder-ONly approach (MADEON) consists of a single decoder that takes speech tokens as a condition and predicts text tokens in an autoregressive manner. To enhance MADEON, we further propose speech prefixing that performs bidirectional processing on speech tokens, which enriches the contextual information in the hidden states. Our experiments show that MADEON significantly outperforms a non-selective SSM. The combination of speech prefixing and the recently proposed Mamba-2 yields comparable performance to Transformer-based models on large datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06968
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
Masuyama, Yoshiki
Miyazaki, Koichi
Murata, Masato
Sound
Audio and Speech Processing
Selective state space models (SSMs) represented by Mamba have demonstrated their computational efficiency and promising outcomes in various tasks, including automatic speech recognition (ASR). Mamba has been applied to ASR task with the attention-based encoder-decoder framework, where the cross-attention mechanism between encoder and decoder remains. This paper explores the capability of Mamba as the decoder-only architecture in ASR task. Our MAmba-based DEcoder-ONly approach (MADEON) consists of a single decoder that takes speech tokens as a condition and predicts text tokens in an autoregressive manner. To enhance MADEON, we further propose speech prefixing that performs bidirectional processing on speech tokens, which enriches the contextual information in the hidden states. Our experiments show that MADEON significantly outperforms a non-selective SSM. The combination of speech prefixing and the recently proposed Mamba-2 yields comparable performance to Transformer-based models on large datasets.
title Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.06968