A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Shenghui, Huang, Hukai, Yao, Jinanglong, Wang, Kaidi, Hong, Qingyang, Li, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918042758283264
author Lu, Shenghui
Huang, Hukai
Yao, Jinanglong
Wang, Kaidi
Hong, Qingyang
Li, Lin
author_facet Lu, Shenghui
Huang, Hukai
Yao, Jinanglong
Wang, Kaidi
Hong, Qingyang
Li, Lin
contents This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins for single-channel speech enhancement. The sub-band module captures surrounding frequency bin information at the input, while the deep filtering module applies filtering at the output to both the target TF bin and its surrounding TF bins. To further improve the model performance, we decouple deep filtering into temporal and frequency components and introduce a two-stage framework, reducing the complexity of filter coefficient prediction at each stage. Additionally, we propose the TAConv module to strengthen convolutional feature extraction. Experimental results demonstrate that the proposed hierarchical deep filtering network (HDF-Net) effectively utilizes surrounding TF bin information and outperforms other advanced systems while using fewer resources.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01023
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
Lu, Shenghui
Huang, Hukai
Yao, Jinanglong
Wang, Kaidi
Hong, Qingyang
Li, Lin
Sound
Artificial Intelligence
Audio and Speech Processing
This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins for single-channel speech enhancement. The sub-band module captures surrounding frequency bin information at the input, while the deep filtering module applies filtering at the output to both the target TF bin and its surrounding TF bins. To further improve the model performance, we decouple deep filtering into temporal and frequency components and introduce a two-stage framework, reducing the complexity of filter coefficient prediction at each stage. Additionally, we propose the TAConv module to strengthen convolutional feature extraction. Experimental results demonstrate that the proposed hierarchical deep filtering network (HDF-Net) effectively utilizes surrounding TF bin information and outperforms other advanced systems while using fewer resources.
title A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.01023