Universal Speech Enhancement with Regression and Generative Mamba

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chao, Rong, Nasretdinov, Rauf, Wang, Yu-Chiang Frank, Jukić, Ante, Fu, Szu-Wei, Tsao, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911184792322048
author Chao, Rong
Nasretdinov, Rauf
Wang, Yu-Chiang Frank
Jukić, Ante
Fu, Szu-Wei
Tsao, Yu
author_facet Chao, Rong
Nasretdinov, Rauf
Wang, Yu-Chiang Frank
Jukić, Ante
Fu, Szu-Wei
Tsao, Yu
contents The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Speech Enhancement with Regression and Generative Mamba
Chao, Rong
Nasretdinov, Rauf
Wang, Yu-Chiang Frank
Jukić, Ante
Fu, Szu-Wei
Tsao, Yu
Sound
Audio and Speech Processing
The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions.
title Universal Speech Enhancement with Regression and Generative Mamba
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.21198