Robust Training for Speaker Verification against Noisy Labels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Zhihua, He, Liang, Ma, Hanhan, Guo, Xiaochen, Li, Lin
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908998089834496
author Fang, Zhihua
He, Liang
Ma, Hanhan
Guo, Xiaochen
Li, Lin
author_facet Fang, Zhihua
He, Liang
Ma, Hanhan
Guo, Xiaochen
Li, Lin
contents The deep learning models used for speaker verification rely heavily on large amounts of data and correct labeling. However, noisy (incorrect) labels often occur, which degrades the performance of the system. In this paper, we propose a novel two-stage learning method to filter out noisy labels from speaker datasets. Since a DNN will first fit data with clean labels, we first train the model with all data for several epochs. Then, based on this model, the model predictions are compared with the labels using our proposed the OR-Gate with top-k mechanism to select the data with clean labels and the selected data is used to train the model. This process is iterated until the training is completed. We have demonstrated the effectiveness of this method in filtering noisy labels through extensive experiments and have achieved excellent performance on the VoxCeleb (1 and 2) with different added noise rates.
format Preprint
id arxiv_https___arxiv_org_abs_2211_12080
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Robust Training for Speaker Verification against Noisy Labels
Fang, Zhihua
He, Liang
Ma, Hanhan
Guo, Xiaochen
Li, Lin
Sound
Audio and Speech Processing
The deep learning models used for speaker verification rely heavily on large amounts of data and correct labeling. However, noisy (incorrect) labels often occur, which degrades the performance of the system. In this paper, we propose a novel two-stage learning method to filter out noisy labels from speaker datasets. Since a DNN will first fit data with clean labels, we first train the model with all data for several epochs. Then, based on this model, the model predictions are compared with the labels using our proposed the OR-Gate with top-k mechanism to select the data with clean labels and the selected data is used to train the model. This process is iterated until the training is completed. We have demonstrated the effectiveness of this method in filtering noisy labels through extensive experiments and have achieved excellent performance on the VoxCeleb (1 and 2) with different added noise rates.
title Robust Training for Speaker Verification against Noisy Labels
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2211.12080