Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916544746881024 |
|---|---|
| author | Wen, Wen Zhou, Qiang Xi, Yu Li, Haoyu Gong, Ziqi Yu, Kai |
| author_facet | Wen, Wen Zhou, Qiang Xi, Yu Li, Haoyu Gong, Ziqi Yu, Kai |
| contents | In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in extremely low signal-to-noise ratio (SNR) conditions. To tackle this issue, we propose a triple-steering spatial selection method, a flexible framework that uses three steering vectors to guide enhancement and determine the enhancement range. Specifically, we introduce a causal-directed U-Net (CDUNet) model, which takes raw multi-channel speech and the desired enhancement width as inputs. This enables dynamic adjustment of steering vectors based on the target direction and fine-tuning of the enhancement region according to the angular separation between the target and interference signals. Our model with only a dual microphone array, excels in both speech quality and downstream task performance. It operates in real-time with minimal parameters, making it ideal for low-latency, on-device streaming applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_18141 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario Wen, Wen Zhou, Qiang Xi, Yu Li, Haoyu Gong, Ziqi Yu, Kai Audio and Speech Processing Sound In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in extremely low signal-to-noise ratio (SNR) conditions. To tackle this issue, we propose a triple-steering spatial selection method, a flexible framework that uses three steering vectors to guide enhancement and determine the enhancement range. Specifically, we introduce a causal-directed U-Net (CDUNet) model, which takes raw multi-channel speech and the desired enhancement width as inputs. This enables dynamic adjustment of steering vectors based on the target direction and fine-tuning of the enhancement region according to the angular separation between the target and interference signals. Our model with only a dual microphone array, excels in both speech quality and downstream task performance. It operates in real-time with minimal parameters, making it ideal for low-latency, on-device streaming applications. |
| title | Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2412.18141 |