Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Labrador, Beltrán, Otero-Gonzalez, Manuel, Lozano-Diez, Alicia, Ramos, Daniel, Toledano, Doroteo T., Gonzalez-Rodriguez, Joaquin
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917569846312960
author Labrador, Beltrán
Otero-Gonzalez, Manuel
Lozano-Diez, Alicia
Ramos, Daniel
Toledano, Doroteo T.
Gonzalez-Rodriguez, Joaquin
author_facet Labrador, Beltrán
Otero-Gonzalez, Manuel
Lozano-Diez, Alicia
Ramos, Daniel
Toledano, Doroteo T.
Gonzalez-Rodriguez, Joaquin
contents This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles, noises, and channel distortions. It includes 160 Spanish celebrities spanning various categories, ensuring a representative distribution across age groups and geographic regions in Spain. We provide two speaker trial lists for speaker identification tasks, each of them with same-video or different-video target trials respectively, accompanied by a cross-lingual evaluation of ResNet pretrained models. Preliminary speaker identification results suggest that the complexity of the detection task in VoxCeleb-ESP is equivalent to that of the original and much larger VoxCeleb in English. VoxCeleb-ESP contributes to the expansion of speaker recognition benchmarks with a comprehensive and diverse dataset for the Spanish language.
format Preprint
id arxiv_https___arxiv_org_abs_2401_09441
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices
Labrador, Beltrán
Otero-Gonzalez, Manuel
Lozano-Diez, Alicia
Ramos, Daniel
Toledano, Doroteo T.
Gonzalez-Rodriguez, Joaquin
Sound
Machine Learning
Audio and Speech Processing
This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles, noises, and channel distortions. It includes 160 Spanish celebrities spanning various categories, ensuring a representative distribution across age groups and geographic regions in Spain. We provide two speaker trial lists for speaker identification tasks, each of them with same-video or different-video target trials respectively, accompanied by a cross-lingual evaluation of ResNet pretrained models. Preliminary speaker identification results suggest that the complexity of the detection task in VoxCeleb-ESP is equivalent to that of the original and much larger VoxCeleb in English. VoxCeleb-ESP contributes to the expansion of speaker recognition benchmarks with a comprehensive and diverse dataset for the Spanish language.
title Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2401.09441