All Neural Low-latency Directional Speech Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911946595368960 |
|---|---|
| author | Pandey, Ashutosh Lee, Sanha Azcarreta, Juan Wong, Daniel Xu, Buye |
| author_facet | Pandey, Ashutosh Lee, Sanha Azcarreta, Juan Wong, Daniel Xu, Buye |
| contents | We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are transformed and fused into a recurrent neural network based speech extraction model. This process enables the model to effectively extract speech from a specified DOA. Unlike previous methods that relied on hand-crafted directional features, the proposed model trains DOA embeddings from scratch using speech enhancement loss, making it suitable for low-latency scenarios. Additionally, it operates at a high frame rate, taking in DOA with each input frame, which brings in the capability of quickly adapting to changing scene in highly dynamic real-world scenarios. We provide extensive evaluation to demonstrate the model's efficacy in directional speech extraction, robustness to DOA mismatch, and its capability to quickly adapt to abrupt changes in DOA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_04879 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | All Neural Low-latency Directional Speech Extraction Pandey, Ashutosh Lee, Sanha Azcarreta, Juan Wong, Daniel Xu, Buye Sound Audio and Speech Processing We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are transformed and fused into a recurrent neural network based speech extraction model. This process enables the model to effectively extract speech from a specified DOA. Unlike previous methods that relied on hand-crafted directional features, the proposed model trains DOA embeddings from scratch using speech enhancement loss, making it suitable for low-latency scenarios. Additionally, it operates at a high frame rate, taking in DOA with each input frame, which brings in the capability of quickly adapting to changing scene in highly dynamic real-world scenarios. We provide extensive evaluation to demonstrate the model's efficacy in directional speech extraction, robustness to DOA mismatch, and its capability to quickly adapt to abrupt changes in DOA. |
| title | All Neural Low-latency Directional Speech Extraction |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2407.04879 |