Schrödinger Bridge Mamba for One-Step Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jing, Wang, Sirui, Wu, Chao, Guo, Lei, Fan, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915836301672448
author Yang, Jing
Wang, Sirui
Wu, Chao
Guo, Lei
Fan, Fan
author_facet Yang, Jing
Wang, Sirui
Wu, Chao
Guo, Lei
Fan, Fan
contents We present Schrödinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schrödinger Bridge (SB) training paradigm and the Mamba architecture. Experiments of joint denoising and dereverberation tasks demonstrate SBM outperforms strong generative and discriminative methods on multiple metrics with only one step of inference while achieving a competitive real-time factor for streaming feasibility. Ablation studies reveal that the SB paradigm consistently yields improved performance across diverse architectures over conventional mapping. Furthermore, Mamba exhibits a stronger performance under the SB paradigm compared to Multi-Head Self-Attention (MHSA) and Long Short-Term Memory (LSTM) backbones. These findings highlight the synergy between the Mamba architecture and the SB trajectory-based training, providing a high-quality solution for real-world speech enhancement. Demo page: https://sbmse.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2510_16834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Schrödinger Bridge Mamba for One-Step Speech Enhancement
Yang, Jing
Wang, Sirui
Wu, Chao
Guo, Lei
Fan, Fan
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
We present Schrödinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schrödinger Bridge (SB) training paradigm and the Mamba architecture. Experiments of joint denoising and dereverberation tasks demonstrate SBM outperforms strong generative and discriminative methods on multiple metrics with only one step of inference while achieving a competitive real-time factor for streaming feasibility. Ablation studies reveal that the SB paradigm consistently yields improved performance across diverse architectures over conventional mapping. Furthermore, Mamba exhibits a stronger performance under the SB paradigm compared to Multi-Head Self-Attention (MHSA) and Long Short-Term Memory (LSTM) backbones. These findings highlight the synergy between the Mamba architecture and the SB trajectory-based training, providing a high-quality solution for real-world speech enhancement. Demo page: https://sbmse.github.io
title Schrödinger Bridge Mamba for One-Step Speech Enhancement
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2510.16834