FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tabatabaee, Saba, Liu, Jing, Espy-Wilson, Carol
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912558468825088
author Tabatabaee, Saba
Liu, Jing
Espy-Wilson, Carol
author_facet Tabatabaee, Saba
Liu, Jing
Espy-Wilson, Carol
contents Creating Speaker Verification (SV) systems for classroom settings that are robust to classroom noises such as babble noise is crucial for the development of AI tools that assist educational environments. In this work, we study the efficacy of finetuning with augmented children datasets to adapt the x-vector and ECAPA-TDNN to classroom environments. We demonstrate that finetuning with augmented children's datasets is powerful in that regard and reduces the Equal Error Rate (EER) of x-vector and ECAPA-TDNN models for both classroom datasets and children speech datasets. Notably, this method reduces EER of the ECAPA-TDNN model on average by half (a 5 % improvement) for classrooms in the MPT dataset compared to the ECAPA-TDNN baseline model. The x-vector model shows an 8 % average improvement for classrooms in the NCTE dataset compared to its baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments
Tabatabaee, Saba
Liu, Jing
Espy-Wilson, Carol
Audio and Speech Processing
Creating Speaker Verification (SV) systems for classroom settings that are robust to classroom noises such as babble noise is crucial for the development of AI tools that assist educational environments. In this work, we study the efficacy of finetuning with augmented children datasets to adapt the x-vector and ECAPA-TDNN to classroom environments. We demonstrate that finetuning with augmented children's datasets is powerful in that regard and reduces the Equal Error Rate (EER) of x-vector and ECAPA-TDNN models for both classroom datasets and children speech datasets. Notably, this method reduces EER of the ECAPA-TDNN model on average by half (a 5 % improvement) for classrooms in the MPT dataset compared to the ECAPA-TDNN baseline model. The x-vector model shows an 8 % average improvement for classrooms in the NCTE dataset compared to its baseline.
title FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments
topic Audio and Speech Processing
url https://arxiv.org/abs/2505.20222