SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Zhongshu, Lei, Tong, Hu, Qinwen, Cao, Zhanzhong, Tang, Ming, Lu, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909289762783232
author Hou, Zhongshu
Lei, Tong
Hu, Qinwen
Cao, Zhanzhong
Tang, Ming
Lu, Jing
author_facet Hou, Zhongshu
Lei, Tong
Hu, Qinwen
Cao, Zhanzhong
Tang, Ming
Lu, Jing
contents Despite significant progress made in the last decade, deep neural network (DNN) based speech enhancement (SE) still faces the challenge of notable degradation in the quality of recovered speech under low signal-to-noise ratio (SNR) conditions. In this letter, we propose an SNR-progressive speech enhancement model with harmonic compensation for low-SNR SE. Reliable pitch estimation is obtained from the intermediate output, which has the benefit of retaining more speech components than the coarse estimate while possessing a significant higher SNR than the input noisy speech. An effective harmonic compensation mechanism is introduced for better harmonic recovery. Extensive ex-periments demonstrate the advantage of our proposed model. A multi-modal speech extraction system based on the proposed backbone model ranks first in the ICASSP 2024 MISP Challenge: https://mispchallenge.github.io/mispchallenge2023/index.html.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16317
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement
Hou, Zhongshu
Lei, Tong
Hu, Qinwen
Cao, Zhanzhong
Tang, Ming
Lu, Jing
Sound
Audio and Speech Processing
Despite significant progress made in the last decade, deep neural network (DNN) based speech enhancement (SE) still faces the challenge of notable degradation in the quality of recovered speech under low signal-to-noise ratio (SNR) conditions. In this letter, we propose an SNR-progressive speech enhancement model with harmonic compensation for low-SNR SE. Reliable pitch estimation is obtained from the intermediate output, which has the benefit of retaining more speech components than the coarse estimate while possessing a significant higher SNR than the input noisy speech. An effective harmonic compensation mechanism is introduced for better harmonic recovery. Extensive ex-periments demonstrate the advantage of our proposed model. A multi-modal speech extraction system based on the proposed backbone model ranks first in the ICASSP 2024 MISP Challenge: https://mispchallenge.github.io/mispchallenge2023/index.html.
title SNR-Progressive Model with Harmonic Compensation for Low-SNR Speech Enhancement
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.16317