SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929408202244096 |
|---|---|
| author | Lin, Jingru Ge, Meng Ao, Junyi Deng, Liqun Li, Haizhou |
| author_facet | Lin, Jingru Ge, Meng Ao, Junyi Deng, Liqun Li, Haizhou |
| contents | It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on single-speaker speech data, limiting their effectiveness in mixture speech. This motivates us to explore pre-training on mixture speech. This work presents SA-WavLM, a novel pre-trained model for mixture speech. Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction. In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers. Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence. Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_02826 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech Lin, Jingru Ge, Meng Ao, Junyi Deng, Liqun Li, Haizhou Audio and Speech Processing It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on single-speaker speech data, limiting their effectiveness in mixture speech. This motivates us to explore pre-training on mixture speech. This work presents SA-WavLM, a novel pre-trained model for mixture speech. Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction. In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers. Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence. Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models. |
| title | SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2407.02826 |