3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Shulin, liu, Jinjiang, Li, Hao, Yang, Yang, Chen, Fei, Zhang, Xueliang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910287688368128
author He, Shulin
liu, Jinjiang
Li, Hao
Yang, Yang
Chen, Fei
Zhang, Xueliang
author_facet He, Shulin
liu, Jinjiang
Li, Hao
Yang, Yang
Chen, Fei
Zhang, Xueliang
contents Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction.
format Preprint
id arxiv_https___arxiv_org_abs_2312_10979
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
He, Shulin
liu, Jinjiang
Li, Hao
Yang, Yang
Chen, Fei
Zhang, Xueliang
Sound
Audio and Speech Processing
Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction.
title 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2312.10979