Mamba in Speech: Towards an Alternative to Self-Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912347982921728 |
|---|---|
| author | Zhang, Xiangyu Zhang, Qiquan Liu, Hexin Xiao, Tianyi Qian, Xinyuan Ahmed, Beena Ambikairajah, Eliathamby Li, Haizhou Epps, Julien |
| author_facet | Zhang, Xiangyu Zhang, Qiquan Liu, Hexin Xiao, Tianyi Qian, Xinyuan Ahmed, Beena Ambikairajah, Eliathamby Li, Haizhou Epps, Julien |
| contents | Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing by discussing two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The experimental results confirm that bidirectional Mamba (BiMamba) consistently outperforms vanilla Mamba, highlighting the advantages of a bidirectional design for speech processing. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in the Transformer model and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section, offering insights for extending this research to a broader scope of tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_12609 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Mamba in Speech: Towards an Alternative to Self-Attention Zhang, Xiangyu Zhang, Qiquan Liu, Hexin Xiao, Tianyi Qian, Xinyuan Ahmed, Beena Ambikairajah, Eliathamby Li, Haizhou Epps, Julien Audio and Speech Processing Sound Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing by discussing two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The experimental results confirm that bidirectional Mamba (BiMamba) consistently outperforms vanilla Mamba, highlighting the advantages of a bidirectional design for speech processing. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in the Transformer model and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section, offering insights for extending this research to a broader scope of tasks. |
| title | Mamba in Speech: Towards an Alternative to Self-Attention |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2405.12609 |