State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barahona, Sara, Mošner, Ladislav, Stafylakis, Themos, Plchot, Oldřich, Peng, Junyi, Burget, Lukáš, Černocký, Jan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912733004300288
author Barahona, Sara
Mošner, Ladislav
Stafylakis, Themos
Plchot, Oldřich
Peng, Junyi
Burget, Lukáš
Černocký, Jan
author_facet Barahona, Sara
Mošner, Ladislav
Stafylakis, Themos
Plchot, Oldřich
Peng, Junyi
Burget, Lukáš
Černocký, Jan
contents In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments belonging to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02364
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
Barahona, Sara
Mošner, Ladislav
Stafylakis, Themos
Plchot, Oldřich
Peng, Junyi
Burget, Lukáš
Černocký, Jan
Audio and Speech Processing
In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments belonging to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation.
title State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
topic Audio and Speech Processing
url https://arxiv.org/abs/2410.02364