Universal Speech Enhancement with Regression and Generative Mamba
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911184792322048 |
|---|---|
| author | Chao, Rong Nasretdinov, Rauf Wang, Yu-Chiang Frank Jukić, Ante Fu, Szu-Wei Tsao, Yu |
| author_facet | Chao, Rong Nasretdinov, Rauf Wang, Yu-Chiang Frank Jukić, Ante Fu, Szu-Wei Tsao, Yu |
| contents | The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_21198 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Universal Speech Enhancement with Regression and Generative Mamba Chao, Rong Nasretdinov, Rauf Wang, Yu-Chiang Frank Jukić, Ante Fu, Szu-Wei Tsao, Yu Sound Audio and Speech Processing The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions. |
| title | Universal Speech Enhancement with Regression and Generative Mamba |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.21198 |