Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dai, Wang, Politis, Archontis, Virtanen, Tuomas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918049628553216
author Dai, Wang
Politis, Archontis
Virtanen, Tuomas
author_facet Dai, Wang
Politis, Archontis
Virtanen, Tuomas
contents We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relative differences, while discrete cues (e.g., language, gender, emotion) retain their categorical distinctions. Compared to fixed speech attribute classification, inter-speaker relative cues offer greater flexibility, facilitating much easier expansion of text-guided target speech extraction datasets. Our experiments show that combining all relative cues yields better performance than random subsets, with gender and temporal order being the most robust across languages and reverberant conditions. Additional cues, such as pitch level, loudness, distance, speaking duration, language, and pitch range, also demonstrate notable benefits in complex scenarios. Fine-tuning pre-trained WavLM Base+ CNN encoders improves overall performance over the Conv1d baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
Dai, Wang
Politis, Archontis
Virtanen, Tuomas
Audio and Speech Processing
Sound
We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order, age, pitch level) are grouped by relative differences, while discrete cues (e.g., language, gender, emotion) retain their categorical distinctions. Compared to fixed speech attribute classification, inter-speaker relative cues offer greater flexibility, facilitating much easier expansion of text-guided target speech extraction datasets. Our experiments show that combining all relative cues yields better performance than random subsets, with gender and temporal order being the most robust across languages and reverberant conditions. Additional cues, such as pitch level, loudness, distance, speaking duration, language, and pitch range, also demonstrate notable benefits in complex scenarios. Fine-tuning pre-trained WavLM Base+ CNN encoders improves overall performance over the Conv1d baseline.
title Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.01483