Experimenting with Additive Margins for Contrastive Self-Supervised Speaker Verification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lepage, Theo, Dehak, Reda
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908417501691904
author Lepage, Theo
Dehak, Reda
author_facet Lepage, Theo
Dehak, Reda
contents Most state-of-the-art self-supervised speaker verification systems rely on a contrastive-based objective function to learn speaker representations from unlabeled speech data. We explore different ways to improve the performance of these methods by: (1) revisiting how positive and negative pairs are sampled through a "symmetric" formulation of the contrastive loss; (2) introducing margins similar to AM-Softmax and AAM-Softmax that have been widely adopted in the supervised setting. We demonstrate the effectiveness of the symmetric contrastive loss which provides more supervision for the self-supervised task. Moreover, we show that Additive Margin and Additive Angular Margin allow reducing the overall number of false negatives and false positives by improving speaker separability. Finally, by combining both techniques and training a larger model we achieve 7.50% EER and 0.5804 minDCF on the VoxCeleb1 test set, which outperforms other contrastive self supervised methods on speaker verification.
format Preprint
id arxiv_https___arxiv_org_abs_2306_03664
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Experimenting with Additive Margins for Contrastive Self-Supervised Speaker Verification
Lepage, Theo
Dehak, Reda
Audio and Speech Processing
Machine Learning
Most state-of-the-art self-supervised speaker verification systems rely on a contrastive-based objective function to learn speaker representations from unlabeled speech data. We explore different ways to improve the performance of these methods by: (1) revisiting how positive and negative pairs are sampled through a "symmetric" formulation of the contrastive loss; (2) introducing margins similar to AM-Softmax and AAM-Softmax that have been widely adopted in the supervised setting. We demonstrate the effectiveness of the symmetric contrastive loss which provides more supervision for the self-supervised task. Moreover, we show that Additive Margin and Additive Angular Margin allow reducing the overall number of false negatives and false positives by improving speaker separability. Finally, by combining both techniques and training a larger model we achieve 7.50% EER and 0.5804 minDCF on the VoxCeleb1 test set, which outperforms other contrastive self supervised methods on speaker verification.
title Experimenting with Additive Margins for Contrastive Self-Supervised Speaker Verification
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2306.03664