Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Wan, Chen, Junhui, Wang, Tianhao, Zhou, Zhenyu, Li, Lantian, Wang, Dong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915370646896640
author Lin, Wan
Chen, Junhui
Wang, Tianhao
Zhou, Zhenyu
Li, Lantian
Wang, Dong
author_facet Lin, Wan
Chen, Junhui
Wang, Tianhao
Zhou, Zhenyu
Li, Lantian
Wang, Dong
contents Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with multi-talker speech due to the unidentifiability of embedding vectors. In this paper, we propose Neural Scoring (NS), a refreshed end-to-end framework that directly estimates verification posterior probabilities without relying on test-side embeddings, making it more robust to complex conditions, e.g., with multiple talkers. To make the training of such an end-to-end model more efficient, we introduce a large-scale trial e2e training (LtE2E) strategy, where each test utterance pairs with a set of enrolled speakers, thus enabling the processing of large-scale verification trials per batch. Experiments on the VoxCeleb dataset demonstrate that NS consistently outperforms both the baseline and competitive methods across various conditions, achieving an overall 70.36% reduction in EER compared to the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16428
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
Lin, Wan
Chen, Junhui
Wang, Tianhao
Zhou, Zhenyu
Li, Lantian
Wang, Dong
Sound
Audio and Speech Processing
Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with multi-talker speech due to the unidentifiability of embedding vectors. In this paper, we propose Neural Scoring (NS), a refreshed end-to-end framework that directly estimates verification posterior probabilities without relying on test-side embeddings, making it more robust to complex conditions, e.g., with multiple talkers. To make the training of such an end-to-end model more efficient, we introduce a large-scale trial e2e training (LtE2E) strategy, where each test utterance pairs with a set of enrolled speakers, thus enabling the processing of large-scale verification trials per batch. Experiments on the VoxCeleb dataset demonstrate that NS consistently outperforms both the baseline and competitive methods across various conditions, achieving an overall 70.36% reduction in EER compared to the baseline.
title Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.16428