Adaptive Convolution for CNN-based Speech Enhancement Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Dahan, Rong, Xiaobin, Sun, Shiruo, Hu, Yuxiang, Zhu, Changbao, Lu, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908639428608000
author Wang, Dahan
Rong, Xiaobin
Sun, Shiruo
Hu, Yuxiang
Zhu, Changbao
Lu, Jing
author_facet Wang, Dahan
Rong, Xiaobin
Sun, Shiruo
Hu, Yuxiang
Zhu, Changbao
Lu, Jing
contents Deep learning-based speech enhancement methods have significantly improved speech quality and intelligibility. Convolutional neural networks (CNNs) have been proven to be essential components of many high-performance models. In this paper, we introduce adaptive convolution, an efficient and versatile convolutional module that enhances the model's capability to adaptively represent speech signals. Adaptive convolution performs frame-wise causal dynamic convolution, generating time-varying kernels for each frame by assembling multiple parallel candidate kernels. A lightweight attention mechanism is proposed for adaptive convolution, leveraging both current and historical information to assign adaptive weights to each candidate kernel. This enables the convolution operation to adapt to frame-level speech spectral features, leading to more efficient extraction and reconstruction. We integrate adaptive convolution into various CNN-based models, highlighting its generalizability. Experimental results demonstrate that adaptive convolution significantly improves the performance with negligible increases in computational complexity, especially for lightweight models. Moreover, we present an intuitive analysis revealing a strong correlation between kernel selection and signal characteristics. Furthermore, we propose the adaptive convolutional recurrent network (AdaptCRN), an ultra-lightweight model that incorporates adaptive convolution and an efficient encoder-decoder design, achieving superior performance compared to models with similar or even higher computational costs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Convolution for CNN-based Speech Enhancement Models
Wang, Dahan
Rong, Xiaobin
Sun, Shiruo
Hu, Yuxiang
Zhu, Changbao
Lu, Jing
Audio and Speech Processing
Sound
Deep learning-based speech enhancement methods have significantly improved speech quality and intelligibility. Convolutional neural networks (CNNs) have been proven to be essential components of many high-performance models. In this paper, we introduce adaptive convolution, an efficient and versatile convolutional module that enhances the model's capability to adaptively represent speech signals. Adaptive convolution performs frame-wise causal dynamic convolution, generating time-varying kernels for each frame by assembling multiple parallel candidate kernels. A lightweight attention mechanism is proposed for adaptive convolution, leveraging both current and historical information to assign adaptive weights to each candidate kernel. This enables the convolution operation to adapt to frame-level speech spectral features, leading to more efficient extraction and reconstruction. We integrate adaptive convolution into various CNN-based models, highlighting its generalizability. Experimental results demonstrate that adaptive convolution significantly improves the performance with negligible increases in computational complexity, especially for lightweight models. Moreover, we present an intuitive analysis revealing a strong correlation between kernel selection and signal characteristics. Furthermore, we propose the adaptive convolutional recurrent network (AdaptCRN), an ultra-lightweight model that incorporates adaptive convolution and an efficient encoder-decoder design, achieving superior performance compared to models with similar or even higher computational costs.
title Adaptive Convolution for CNN-based Speech Enhancement Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2502.14224