Challenging margin-based speaker embedding extractors by using the variational information bottleneck

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stafylakis, Themos, Silnova, Anna, Rohdin, Johan, Plchot, Oldrich, Burget, Lukas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913396824211456
author Stafylakis, Themos
Silnova, Anna
Rohdin, Johan
Plchot, Oldrich
Burget, Lukas
author_facet Stafylakis, Themos
Silnova, Anna
Rohdin, Johan
Plchot, Oldrich
Burget, Lukas
contents Speaker embedding extractors are typically trained using a classification loss over the training speakers. During the last few years, the standard softmax/cross-entropy loss has been replaced by the margin-based losses, yielding significant improvements in speaker recognition accuracy. Motivated by the fact that the margin merely reduces the logit of the target speaker during training, we consider a probabilistic framework that has a similar effect. The variational information bottleneck provides a principled mechanism for making deterministic nodes stochastic, resulting in an implicit reduction of the posterior of the target speaker. We experiment with a wide range of speaker recognition benchmarks and scoring methods and report competitive results to those obtained with the state-of-the-art Additive Angular Margin loss.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12622
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Challenging margin-based speaker embedding extractors by using the variational information bottleneck
Stafylakis, Themos
Silnova, Anna
Rohdin, Johan
Plchot, Oldrich
Burget, Lukas
Audio and Speech Processing
Speaker embedding extractors are typically trained using a classification loss over the training speakers. During the last few years, the standard softmax/cross-entropy loss has been replaced by the margin-based losses, yielding significant improvements in speaker recognition accuracy. Motivated by the fact that the margin merely reduces the logit of the target speaker during training, we consider a probabilistic framework that has a similar effect. The variational information bottleneck provides a principled mechanism for making deterministic nodes stochastic, resulting in an implicit reduction of the posterior of the target speaker. We experiment with a wide range of speaker recognition benchmarks and scoring methods and report competitive results to those obtained with the state-of-the-art Additive Angular Margin loss.
title Challenging margin-based speaker embedding extractors by using the variational information bottleneck
topic Audio and Speech Processing
url https://arxiv.org/abs/2406.12622