SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Jingru, Ge, Meng, Ao, Junyi, Deng, Liqun, Li, Haizhou
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929408202244096
author Lin, Jingru
Ge, Meng
Ao, Junyi
Deng, Liqun
Li, Haizhou
author_facet Lin, Jingru
Ge, Meng
Ao, Junyi
Deng, Liqun
Li, Haizhou
contents It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on single-speaker speech data, limiting their effectiveness in mixture speech. This motivates us to explore pre-training on mixture speech. This work presents SA-WavLM, a novel pre-trained model for mixture speech. Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction. In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers. Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence. Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02826
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
Lin, Jingru
Ge, Meng
Ao, Junyi
Deng, Liqun
Li, Haizhou
Audio and Speech Processing
It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on single-speaker speech data, limiting their effectiveness in mixture speech. This motivates us to explore pre-training on mixture speech. This work presents SA-WavLM, a novel pre-trained model for mixture speech. Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction. In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers. Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence. Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models.
title SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
topic Audio and Speech Processing
url https://arxiv.org/abs/2407.02826