Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917569846312960 |
|---|---|
| author | Labrador, Beltrán Otero-Gonzalez, Manuel Lozano-Diez, Alicia Ramos, Daniel Toledano, Doroteo T. Gonzalez-Rodriguez, Joaquin |
| author_facet | Labrador, Beltrán Otero-Gonzalez, Manuel Lozano-Diez, Alicia Ramos, Daniel Toledano, Doroteo T. Gonzalez-Rodriguez, Joaquin |
| contents | This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles, noises, and channel distortions. It includes 160 Spanish celebrities spanning various categories, ensuring a representative distribution across age groups and geographic regions in Spain. We provide two speaker trial lists for speaker identification tasks, each of them with same-video or different-video target trials respectively, accompanied by a cross-lingual evaluation of ResNet pretrained models. Preliminary speaker identification results suggest that the complexity of the detection task in VoxCeleb-ESP is equivalent to that of the original and much larger VoxCeleb in English. VoxCeleb-ESP contributes to the expansion of speaker recognition benchmarks with a comprehensive and diverse dataset for the Spanish language. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_09441 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices Labrador, Beltrán Otero-Gonzalez, Manuel Lozano-Diez, Alicia Ramos, Daniel Toledano, Doroteo T. Gonzalez-Rodriguez, Joaquin Sound Machine Learning Audio and Speech Processing This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles, noises, and channel distortions. It includes 160 Spanish celebrities spanning various categories, ensuring a representative distribution across age groups and geographic regions in Spain. We provide two speaker trial lists for speaker identification tasks, each of them with same-video or different-video target trials respectively, accompanied by a cross-lingual evaluation of ResNet pretrained models. Preliminary speaker identification results suggest that the complexity of the detection task in VoxCeleb-ESP is equivalent to that of the original and much larger VoxCeleb in English. VoxCeleb-ESP contributes to the expansion of speaker recognition benchmarks with a comprehensive and diverse dataset for the Spanish language. |
| title | Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2401.09441 |