Mamba for Streaming ASR Combined with Unimodal Aggregation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fang, Ying, Li, Xiaofei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917879464591360
author Fang, Ying
Li, Xiaofei
author_facet Fang, Ying
Li, Xiaofei
contents This paper works on streaming automatic speech recognition (ASR). Mamba, a recently proposed state space model, has demonstrated the ability to match or surpass Transformers in various tasks while benefiting from a linear complexity advantage. We explore the efficiency of Mamba encoder for streaming ASR and propose an associated lookahead mechanism for leveraging controllable future information. Additionally, a streaming-style unimodal aggregation (UMA) method is implemented, which automatically detects token activity and streamingly triggers token output, and meanwhile aggregates feature frames for better learning token representation. Based on UMA, an early termination (ET) method is proposed to further reduce recognition latency. Experiments conducted on two Mandarin Chinese datasets demonstrate that the proposed model achieves competitive ASR performance in terms of both recognition accuracy and latency.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00070
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mamba for Streaming ASR Combined with Unimodal Aggregation
Fang, Ying
Li, Xiaofei
Audio and Speech Processing
Computation and Language
Sound
This paper works on streaming automatic speech recognition (ASR). Mamba, a recently proposed state space model, has demonstrated the ability to match or surpass Transformers in various tasks while benefiting from a linear complexity advantage. We explore the efficiency of Mamba encoder for streaming ASR and propose an associated lookahead mechanism for leveraging controllable future information. Additionally, a streaming-style unimodal aggregation (UMA) method is implemented, which automatically detects token activity and streamingly triggers token output, and meanwhile aggregates feature frames for better learning token representation. Based on UMA, an early termination (ET) method is proposed to further reduce recognition latency. Experiments conducted on two Mandarin Chinese datasets demonstrate that the proposed model achieves competitive ASR performance in terms of both recognition accuracy and latency.
title Mamba for Streaming ASR Combined with Unimodal Aggregation
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2410.00070