Saved in:
Bibliographic Details
Main Authors: Groot, Sjoerd, Chen, Qinyu, van Gemert, Jan C., Gao, Chang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.11062
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915368148140032
author Groot, Sjoerd
Chen, Qinyu
van Gemert, Jan C.
Gao, Chang
author_facet Groot, Sjoerd
Chen, Qinyu
van Gemert, Jan C.
Gao, Chang
contents This paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https://github.com/lab-emi/CleanUMamba
format Preprint
id arxiv_https___arxiv_org_abs_2410_11062
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
Groot, Sjoerd
Chen, Qinyu
van Gemert, Jan C.
Gao, Chang
Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Audio and Speech Processing
This paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https://github.com/lab-emi/CleanUMamba
title CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
topic Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2410.11062