Selective State Space Model for Monaural Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Moran, Zhang, Qiquan, Wang, Mingjiang, Zhang, Xiangyu, Liu, Hexin, Ambikairaiah, Eliathamby, Chen, Deying
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917832951857152
author Chen, Moran
Zhang, Qiquan
Wang, Mingjiang
Zhang, Xiangyu
Liu, Hexin
Ambikairaiah, Eliathamby
Chen, Deying
author_facet Chen, Moran
Zhang, Qiquan
Wang, Mingjiang
Zhang, Xiangyu
Liu, Hexin
Ambikairaiah, Eliathamby
Chen, Deying
contents Voice user interfaces (VUIs) have facilitated the efficient interactions between humans and machines through spoken commands. Since real-word acoustic scenes are complex, speech enhancement plays a critical role for robust VUI. Transformer and its variants, such as Conformer, have demonstrated cutting-edge results in speech enhancement. However, both of them suffers from the quadratic computational complexity with respect to the sequence length, which hampers their ability to handle long sequences. Recently a novel State Space Model called Mamba, which shows strong capability to handle long sequences with linear complexity, offers a solution to address this challenge. In this paper, we propose a novel hybrid convolution-Mamba backbone, denoted as MambaDC, for speech enhancement. Our MambaDC marries the benefits of convolutional networks to model the local interactions and Mamba's ability for modeling long-range global dependencies. We conduct comprehensive experiments within both basic and state-of-the-art (SoTA) speech enhancement frameworks, on two commonly used training targets. The results demonstrate that MambaDC outperforms Transformer, Conformer, and the standard Mamba across all training targets. Built upon the current advanced framework, the use of MambaDC backbone showcases superior results compared to existing \textcolor{black}{SoTA} systems. This sets the stage for efficient long-range global modeling in speech enhancement.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06217
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Selective State Space Model for Monaural Speech Enhancement
Chen, Moran
Zhang, Qiquan
Wang, Mingjiang
Zhang, Xiangyu
Liu, Hexin
Ambikairaiah, Eliathamby
Chen, Deying
Audio and Speech Processing
Voice user interfaces (VUIs) have facilitated the efficient interactions between humans and machines through spoken commands. Since real-word acoustic scenes are complex, speech enhancement plays a critical role for robust VUI. Transformer and its variants, such as Conformer, have demonstrated cutting-edge results in speech enhancement. However, both of them suffers from the quadratic computational complexity with respect to the sequence length, which hampers their ability to handle long sequences. Recently a novel State Space Model called Mamba, which shows strong capability to handle long sequences with linear complexity, offers a solution to address this challenge. In this paper, we propose a novel hybrid convolution-Mamba backbone, denoted as MambaDC, for speech enhancement. Our MambaDC marries the benefits of convolutional networks to model the local interactions and Mamba's ability for modeling long-range global dependencies. We conduct comprehensive experiments within both basic and state-of-the-art (SoTA) speech enhancement frameworks, on two commonly used training targets. The results demonstrate that MambaDC outperforms Transformer, Conformer, and the standard Mamba across all training targets. Built upon the current advanced framework, the use of MambaDC backbone showcases superior results compared to existing \textcolor{black}{SoTA} systems. This sets the stage for efficient long-range global modeling in speech enhancement.
title Selective State Space Model for Monaural Speech Enhancement
topic Audio and Speech Processing
url https://arxiv.org/abs/2411.06217