Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rostami, Amir Mohammad, Jafarzadeh, Pourya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916012792741888
author Rostami, Amir Mohammad
Jafarzadeh, Pourya
author_facet Rostami, Amir Mohammad
Jafarzadeh, Pourya
contents This paper presents a system for the 2024 Text-Dependent Speaker Verification (TdSV) Challenge. The system achieved a Minimum Detection Cost Function (MinDCF) of 0.0461 and an Equal Error Rate (EER) of 1.3\%. Our approach focused on adapting existing state-of-the-art neural networks, ResNet-TDNN and NeXt-TDNN, originally trained on the VoxCeleb dataset. This strategy was chosen because of the limited challenge duration and the available resources at the time. In addition, we designed a lightweight and resource-efficient model, EfficientNet-A0, trained specifically on the challenge dataset to improve adaptation and strengthen the ensemble approach. Our system combines advanced neural architectures, extensive data augmentation, and optimised hyperparameters. These components helped achieve strong performance in text-dependent speaker verification. The results also demonstrate the effectiveness of multi-model ensemble learning for both speaker and phrase verification.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14896
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report
Rostami, Amir Mohammad
Jafarzadeh, Pourya
Sound
Machine Learning
This paper presents a system for the 2024 Text-Dependent Speaker Verification (TdSV) Challenge. The system achieved a Minimum Detection Cost Function (MinDCF) of 0.0461 and an Equal Error Rate (EER) of 1.3\%. Our approach focused on adapting existing state-of-the-art neural networks, ResNet-TDNN and NeXt-TDNN, originally trained on the VoxCeleb dataset. This strategy was chosen because of the limited challenge duration and the available resources at the time. In addition, we designed a lightweight and resource-efficient model, EfficientNet-A0, trained specifically on the challenge dataset to improve adaptation and strengthen the ensemble approach. Our system combines advanced neural architectures, extensive data augmentation, and optimised hyperparameters. These components helped achieve strong performance in text-dependent speaker verification. The results also demonstrate the effectiveness of multi-model ensemble learning for both speaker and phrase verification.
title Text-Dependent Speaker Verification (TdSV) Challenge 2024: Team Naive System Report
topic Sound
Machine Learning
url https://arxiv.org/abs/2605.14896