Saved in:
Bibliographic Details
Main Authors: Liu, Zikai, Wang, Ziqian, Li, Xingchen, Zhu, Yike, Wang, Shuai, Xiao, Longshuai, Xie, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.06810
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915926055583744
author Liu, Zikai
Wang, Ziqian
Li, Xingchen
Zhu, Yike
Wang, Shuai
Xiao, Longshuai
Xie, Lei
author_facet Liu, Zikai
Wang, Ziqian
Li, Xingchen
Zhu, Yike
Wang, Shuai
Xiao, Longshuai
Xie, Lei
contents Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity of blind source separation, it remains vulnerable to speaker confusion, where models mistakenly extract the interfering speaker. Furthermore, conventional TSE relies on static inference pipeline, where performance is limited by the quality of the fixed enrollment. To overcome these limitations, we propose EvoTSE, an evolving TSE framework in which the enrollment is continuously updated through reliability-filtered retrieval over high-confidence historical estimates. This mechanism reduces speaker confusion and relaxes the quality requirements for pre-recorded enrollment without relying on additional annotated data. Experiments across multiple benchmarks demonstrate that EvoTSE achieves consistent improvements, especially when evaluated on out-of-domain (OOD) scenarios. Our code and checkpoints are available.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06810
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EvoTSE: Evolving Enrollment for Target Speaker Extraction
Liu, Zikai
Wang, Ziqian
Li, Xingchen
Zhu, Yike
Wang, Shuai
Xiao, Longshuai
Xie, Lei
Audio and Speech Processing
Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity of blind source separation, it remains vulnerable to speaker confusion, where models mistakenly extract the interfering speaker. Furthermore, conventional TSE relies on static inference pipeline, where performance is limited by the quality of the fixed enrollment. To overcome these limitations, we propose EvoTSE, an evolving TSE framework in which the enrollment is continuously updated through reliability-filtered retrieval over high-confidence historical estimates. This mechanism reduces speaker confusion and relaxes the quality requirements for pre-recorded enrollment without relying on additional annotated data. Experiments across multiple benchmarks demonstrate that EvoTSE achieves consistent improvements, especially when evaluated on out-of-domain (OOD) scenarios. Our code and checkpoints are available.
title EvoTSE: Evolving Enrollment for Target Speaker Extraction
topic Audio and Speech Processing
url https://arxiv.org/abs/2604.06810