State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912733004300288 |
|---|---|
| author | Barahona, Sara Mošner, Ladislav Stafylakis, Themos Plchot, Oldřich Peng, Junyi Burget, Lukáš Černocký, Jan |
| author_facet | Barahona, Sara Mošner, Ladislav Stafylakis, Themos Plchot, Oldřich Peng, Junyi Burget, Lukáš Černocký, Jan |
| contents | In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments belonging to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02364 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data Barahona, Sara Mošner, Ladislav Stafylakis, Themos Plchot, Oldřich Peng, Junyi Burget, Lukáš Černocký, Jan Audio and Speech Processing In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments belonging to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation. |
| title | State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.02364 |