All Neural Low-latency Directional Speech Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pandey, Ashutosh, Lee, Sanha, Azcarreta, Juan, Wong, Daniel, Xu, Buye
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911946595368960
author Pandey, Ashutosh
Lee, Sanha
Azcarreta, Juan
Wong, Daniel
Xu, Buye
author_facet Pandey, Ashutosh
Lee, Sanha
Azcarreta, Juan
Wong, Daniel
Xu, Buye
contents We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are transformed and fused into a recurrent neural network based speech extraction model. This process enables the model to effectively extract speech from a specified DOA. Unlike previous methods that relied on hand-crafted directional features, the proposed model trains DOA embeddings from scratch using speech enhancement loss, making it suitable for low-latency scenarios. Additionally, it operates at a high frame rate, taking in DOA with each input frame, which brings in the capability of quickly adapting to changing scene in highly dynamic real-world scenarios. We provide extensive evaluation to demonstrate the model's efficacy in directional speech extraction, robustness to DOA mismatch, and its capability to quickly adapt to abrupt changes in DOA.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04879
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle All Neural Low-latency Directional Speech Extraction
Pandey, Ashutosh
Lee, Sanha
Azcarreta, Juan
Wong, Daniel
Xu, Buye
Sound
Audio and Speech Processing
We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are transformed and fused into a recurrent neural network based speech extraction model. This process enables the model to effectively extract speech from a specified DOA. Unlike previous methods that relied on hand-crafted directional features, the proposed model trains DOA embeddings from scratch using speech enhancement loss, making it suitable for low-latency scenarios. Additionally, it operates at a high frame rate, taking in DOA with each input frame, which brings in the capability of quickly adapting to changing scene in highly dynamic real-world scenarios. We provide extensive evaluation to demonstrate the model's efficacy in directional speech extraction, robustness to DOA mismatch, and its capability to quickly adapt to abrupt changes in DOA.
title All Neural Low-latency Directional Speech Extraction
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.04879