Challenging margin-based speaker embedding extractors by using the variational information bottleneck
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913396824211456 |
|---|---|
| author | Stafylakis, Themos Silnova, Anna Rohdin, Johan Plchot, Oldrich Burget, Lukas |
| author_facet | Stafylakis, Themos Silnova, Anna Rohdin, Johan Plchot, Oldrich Burget, Lukas |
| contents | Speaker embedding extractors are typically trained using a classification loss over the training speakers. During the last few years, the standard softmax/cross-entropy loss has been replaced by the margin-based losses, yielding significant improvements in speaker recognition accuracy. Motivated by the fact that the margin merely reduces the logit of the target speaker during training, we consider a probabilistic framework that has a similar effect. The variational information bottleneck provides a principled mechanism for making deterministic nodes stochastic, resulting in an implicit reduction of the posterior of the target speaker. We experiment with a wide range of speaker recognition benchmarks and scoring methods and report competitive results to those obtained with the state-of-the-art Additive Angular Margin loss. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_12622 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Challenging margin-based speaker embedding extractors by using the variational information bottleneck Stafylakis, Themos Silnova, Anna Rohdin, Johan Plchot, Oldrich Burget, Lukas Audio and Speech Processing Speaker embedding extractors are typically trained using a classification loss over the training speakers. During the last few years, the standard softmax/cross-entropy loss has been replaced by the margin-based losses, yielding significant improvements in speaker recognition accuracy. Motivated by the fact that the margin merely reduces the logit of the target speaker during training, we consider a probabilistic framework that has a similar effect. The variational information bottleneck provides a principled mechanism for making deterministic nodes stochastic, resulting in an implicit reduction of the posterior of the target speaker. We experiment with a wide range of speaker recognition benchmarks and scoring methods and report competitive results to those obtained with the state-of-the-art Additive Angular Margin loss. |
| title | Challenging margin-based speaker embedding extractors by using the variational information bottleneck |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.12622 |