Mamba in Speech: Towards an Alternative to Self-Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xiangyu, Zhang, Qiquan, Liu, Hexin, Xiao, Tianyi, Qian, Xinyuan, Ahmed, Beena, Ambikairajah, Eliathamby, Li, Haizhou, Epps, Julien
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912347982921728
author Zhang, Xiangyu
Zhang, Qiquan
Liu, Hexin
Xiao, Tianyi
Qian, Xinyuan
Ahmed, Beena
Ambikairajah, Eliathamby
Li, Haizhou
Epps, Julien
author_facet Zhang, Xiangyu
Zhang, Qiquan
Liu, Hexin
Xiao, Tianyi
Qian, Xinyuan
Ahmed, Beena
Ambikairajah, Eliathamby
Li, Haizhou
Epps, Julien
contents Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing by discussing two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The experimental results confirm that bidirectional Mamba (BiMamba) consistently outperforms vanilla Mamba, highlighting the advantages of a bidirectional design for speech processing. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in the Transformer model and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section, offering insights for extending this research to a broader scope of tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12609
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mamba in Speech: Towards an Alternative to Self-Attention
Zhang, Xiangyu
Zhang, Qiquan
Liu, Hexin
Xiao, Tianyi
Qian, Xinyuan
Ahmed, Beena
Ambikairajah, Eliathamby
Li, Haizhou
Epps, Julien
Audio and Speech Processing
Sound
Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing by discussing two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The experimental results confirm that bidirectional Mamba (BiMamba) consistently outperforms vanilla Mamba, highlighting the advantages of a bidirectional design for speech processing. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in the Transformer model and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section, offering insights for extending this research to a broader scope of tasks.
title Mamba in Speech: Towards an Alternative to Self-Attention
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2405.12609