SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Yongjoon, Choi, Jung-Woo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915856924016640
author Lee, Yongjoon
Choi, Jung-Woo
author_facet Lee, Yongjoon
Choi, Jung-Woo
contents General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose Frequency GLP, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11669
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
Lee, Yongjoon
Choi, Jung-Woo
Audio and Speech Processing
Sound
General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose Frequency GLP, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.
title SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2603.11669