Listen to Extract: Onset-Prompted Target Speaker Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Pengjie, Chen, Kangrui, He, Shulin, Chen, Pengru, Yuan, Shuqi, Kong, He, Zhang, Xueliang, Wang, Zhong-Qiu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915598071496704
author Shen, Pengjie
Chen, Kangrui
He, Shulin
Chen, Pengru
Yuan, Shuqi
Kong, He
Zhang, Xueliang
Wang, Zhong-Qiu
author_facet Shen, Pengjie
Chen, Kangrui
He, Shulin
Chen, Pengru
Yuan, Shuqi
Kong, He
Zhang, Xueliang
Wang, Zhong-Qiu
contents We propose listen to extract (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the speaker's mixed speech with other speakers. For each mixture, LExt concatenates an enrollment utterance of the target speaker to the mixture signal at the waveform level, and trains deep neural networks (DNN) to extract the target speech based on the concatenated mixture signal. The rationale is that, this way, an artificial speech onset is created for the target speaker and it could prompt the DNN (a) which speaker is the target to extract; and (b) spectral-temporal patterns of the target speaker that could help extraction. This simple approach produces strong TSE performance on multiple public TSE datasets including WSJ0-2mix, WHAM! and WHAMR!.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05114
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listen to Extract: Onset-Prompted Target Speaker Extraction
Shen, Pengjie
Chen, Kangrui
He, Shulin
Chen, Pengru
Yuan, Shuqi
Kong, He
Zhang, Xueliang
Wang, Zhong-Qiu
Audio and Speech Processing
Sound
We propose listen to extract (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the speaker's mixed speech with other speakers. For each mixture, LExt concatenates an enrollment utterance of the target speaker to the mixture signal at the waveform level, and trains deep neural networks (DNN) to extract the target speech based on the concatenated mixture signal. The rationale is that, this way, an artificial speech onset is created for the target speaker and it could prompt the DNN (a) which speaker is the target to extract; and (b) spectral-temporal patterns of the target speaker that could help extraction. This simple approach produces strong TSE performance on multiple public TSE datasets including WSJ0-2mix, WHAM! and WHAMR!.
title Listen to Extract: Onset-Prompted Target Speaker Extraction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.05114