pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Ziyang, Lei, Jiahe, Chen, Xueyan, Zhang, Yifan, Pan, Zexu, Xue, Wei, Qian, Xinyuan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917435039285248
author Jiang, Ziyang
Lei, Jiahe
Chen, Xueyan
Zhang, Yifan
Pan, Zexu
Xue, Wei
Qian, Xinyuan
author_facet Jiang, Ziyang
Lei, Jiahe
Chen, Xueyan
Zhang, Yifan
Pan, Zexu
Xue, Wei
Qian, Xinyuan
contents Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded speech, visual information, and spatial information, the acquisition and selection of such strong cues are infeasible in many practical scenarios. Differently, in this paper, we condition the TSE algorithm on semantic cues extracted from limited and unaligned text contents, such as condensed points from a presentation slide. This method is particularly useful in scenarios like meetings, poster sessions, or lecture presentations, where acquiring other cues in real time may be challenging. To this end, we design two different networks. Specifically, our proposed Text Prompt Extractor Network (TPE) fuses audio features with content-based semantic cues to facilitate time-frequency mask generation to filter out extraneous noise. The experimental results show the efficacy in accurately extracting the target speaker's speech by utilizing semantic cues derived from limited and unaligned text, resulting in SI-SDRi of 12.16 dB, SDRi of 12.66 dB, PESQi of 0.830 and STOIi of 0.150.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03109
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
Jiang, Ziyang
Lei, Jiahe
Chen, Xueyan
Zhang, Yifan
Pan, Zexu
Xue, Wei
Qian, Xinyuan
Sound
Multimedia
Audio and Speech Processing
Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded speech, visual information, and spatial information, the acquisition and selection of such strong cues are infeasible in many practical scenarios. Differently, in this paper, we condition the TSE algorithm on semantic cues extracted from limited and unaligned text contents, such as condensed points from a presentation slide. This method is particularly useful in scenarios like meetings, poster sessions, or lecture presentations, where acquiring other cues in real time may be challenging. To this end, we design two different networks. Specifically, our proposed Text Prompt Extractor Network (TPE) fuses audio features with content-based semantic cues to facilitate time-frequency mask generation to filter out extraneous noise. The experimental results show the efficacy in accurately extracting the target speaker's speech by utilizing semantic cues derived from limited and unaligned text, resulting in SI-SDRi of 12.16 dB, SDRi of 12.66 dB, PESQi of 0.830 and STOIi of 0.150.
title pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2411.03109