Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lepage, Théo, Dehak, Réda
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908417459748864
author Lepage, Théo
Dehak, Réda
author_facet Lepage, Théo
Dehak, Réda
contents State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to the amount of data available today. In this study, we explore self-supervised learning for speaker verification by learning representations directly from raw audio. The objective is to produce robust speaker embeddings that have small intra-speaker and large inter-speaker variance. Our approach is based on recent information maximization learning frameworks and an intensive data augmentation pre-processing step. We evaluate the ability of these methods to work without contrastive samples before showing that they achieve better performance when combined with a contrastive loss. Furthermore, we conduct experiments to show that our method reaches competitive results compared to existing techniques and can get better performances compared to a supervised baseline when fine-tuned with a small portion of labeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2207_05506
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
Lepage, Théo
Dehak, Réda
Audio and Speech Processing
Machine Learning
Sound
State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to the amount of data available today. In this study, we explore self-supervised learning for speaker verification by learning representations directly from raw audio. The objective is to produce robust speaker embeddings that have small intra-speaker and large inter-speaker variance. Our approach is based on recent information maximization learning frameworks and an intensive data augmentation pre-processing step. We evaluate the ability of these methods to work without contrastive samples before showing that they achieve better performance when combined with a contrastive loss. Furthermore, we conduct experiments to show that our method reaches competitive results compared to existing techniques and can get better performances compared to a supervised baseline when fine-tuned with a small portion of labeled data.
title Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2207.05506