The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Molavi, Mohammadreza, Khodadadi, Reza
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910713208897536
author Molavi, Mohammadreza
Khodadadi, Reza
author_facet Molavi, Mohammadreza
Khodadadi, Reza
contents This paper introduces an efficient and accurate pipeline for text-dependent speaker verification (TDSV), designed to address the need for high-performance biometric systems. The proposed system incorporates a Fast-Conformer-based ASR module to validate speech content, filtering out Target-Wrong (TW) and Impostor-Wrong (IW) trials. For speaker verification, we propose a feature fusion approach that combines speaker embeddings extracted from wav2vec-BERT and ReDimNet models to create a unified speaker representation. This system achieves competitive results on the TDSV 2024 Challenge test set, with a normalized min-DCF of 0.0452 (rank 2), highlighting its effectiveness in balancing accuracy and robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16276
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024
Molavi, Mohammadreza
Khodadadi, Reza
Sound
Artificial Intelligence
Audio and Speech Processing
This paper introduces an efficient and accurate pipeline for text-dependent speaker verification (TDSV), designed to address the need for high-performance biometric systems. The proposed system incorporates a Fast-Conformer-based ASR module to validate speech content, filtering out Target-Wrong (TW) and Impostor-Wrong (IW) trials. For speaker verification, we propose a feature fusion approach that combines speaker embeddings extracted from wav2vec-BERT and ReDimNet models to create a unified speaker representation. This system achieves competitive results on the TDSV 2024 Challenge test set, with a normalized min-DCF of 0.0452 (rank 2), highlighting its effectiveness in balancing accuracy and robustness.
title The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2411.16276