Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen, Wen, Zhou, Qiang, Xi, Yu, Li, Haoyu, Gong, Ziqi, Yu, Kai
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916544746881024
author Wen, Wen
Zhou, Qiang
Xi, Yu
Li, Haoyu
Gong, Ziqi
Yu, Kai
author_facet Wen, Wen
Zhou, Qiang
Xi, Yu
Li, Haoyu
Gong, Ziqi
Yu, Kai
contents In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in extremely low signal-to-noise ratio (SNR) conditions. To tackle this issue, we propose a triple-steering spatial selection method, a flexible framework that uses three steering vectors to guide enhancement and determine the enhancement range. Specifically, we introduce a causal-directed U-Net (CDUNet) model, which takes raw multi-channel speech and the desired enhancement width as inputs. This enables dynamic adjustment of steering vectors based on the target direction and fine-tuning of the enhancement region according to the angular separation between the target and interference signals. Our model with only a dual microphone array, excels in both speech quality and downstream task performance. It operates in real-time with minimal parameters, making it ideal for low-latency, on-device streaming applications.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18141
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
Wen, Wen
Zhou, Qiang
Xi, Yu
Li, Haoyu
Gong, Ziqi
Yu, Kai
Audio and Speech Processing
Sound
In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in extremely low signal-to-noise ratio (SNR) conditions. To tackle this issue, we propose a triple-steering spatial selection method, a flexible framework that uses three steering vectors to guide enhancement and determine the enhancement range. Specifically, we introduce a causal-directed U-Net (CDUNet) model, which takes raw multi-channel speech and the desired enhancement width as inputs. This enables dynamic adjustment of steering vectors based on the target direction and fine-tuning of the enhancement region according to the angular separation between the target and interference signals. Our model with only a dual microphone array, excels in both speech quality and downstream task performance. It operates in real-time with minimal parameters, making it ideal for low-latency, on-device streaming applications.
title Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2412.18141