CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Farhadipour, Aref, Liu, Shiran, Chapariniya, Masoumeh, Vyshnevetska, Valeriia, Madikeri, Srikanth, Vukovic, Teodora, Dellwo, Volker
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914078726815744
author Farhadipour, Aref
Liu, Shiran
Chapariniya, Masoumeh
Vyshnevetska, Valeriia
Madikeri, Srikanth
Vukovic, Teodora
Dellwo, Volker
author_facet Farhadipour, Aref
Liu, Shiran
Chapariniya, Masoumeh
Vyshnevetska, Valeriia
Madikeri, Srikanth
Vukovic, Teodora
Dellwo, Volker
contents The CL-UZH team submitted one system each for the fixed and open conditions of the NIST SRE 2024 challenge. For the closed-set condition, results for the audio-only trials were achieved using the X-vector system developed with Kaldi. For the audio-visual results we used only models developed for the visual modality. Two sets of results were submitted for the open-set and closed-set conditions, one based on a pretrained model using the VoxBlink2 and VoxCeleb2 datasets. An Xvector-based model was trained from scratch using the CTS superset dataset for the closed set. In addition to the submission of the results of the SRE24 evaluation to the competition website, we talked about the performance of the proposed systems on the SRE24 evaluation in this report.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
Farhadipour, Aref
Liu, Shiran
Chapariniya, Masoumeh
Vyshnevetska, Valeriia
Madikeri, Srikanth
Vukovic, Teodora
Dellwo, Volker
Audio and Speech Processing
Sound
The CL-UZH team submitted one system each for the fixed and open conditions of the NIST SRE 2024 challenge. For the closed-set condition, results for the audio-only trials were achieved using the X-vector system developed with Kaldi. For the audio-visual results we used only models developed for the visual modality. Two sets of results were submitted for the open-set and closed-set conditions, one based on a pretrained model using the VoxBlink2 and VoxCeleb2 datasets. An Xvector-based model was trained from scratch using the CTS superset dataset for the closed set. In addition to the submission of the results of the SRE24 evaluation to the competition website, we talked about the performance of the proposed systems on the SRE24 evaluation in this report.
title CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.00952