Xi+: Uncertainty Supervision for Robust Speaker Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Junjie, Lee, Kong Aik, Truong, Duc-Tuan, Liu, Tianchi, Mak, Man-Wai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911390984306688
author Li, Junjie
Lee, Kong Aik
Truong, Duc-Tuan
Liu, Tianchi
Mak, Man-Wai
author_facet Li, Junjie
Lee, Kong Aik
Truong, Duc-Tuan
Liu, Tianchi
Mak, Man-Wai
contents There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Since individual speech frames do not contribute equally to the utterance-level representation, it is essential to estimate the importance or reliability of each frame. The xi-vector model addresses this by assigning different weights to frames based on uncertainty estimation. However, its uncertainty estimation model is implicitly trained through classification loss alone and does not consider the temporal relationships between frames, which may lead to suboptimal supervision. In this paper, we propose an improved architecture, xi+. Compared to xi-vector, xi+ incorporates a temporal attention module to capture frame-level uncertainty in a context-aware manner. In addition, we introduce a novel loss function, Stochastic Variance Loss, which explicitly supervises the learning of uncertainty. Results demonstrate consistent performance improvements of about 10\% on the VoxCeleb1-O set and 11\% on the NIST SRE 2024 evaluation set.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05993
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Xi+: Uncertainty Supervision for Robust Speaker Embedding
Li, Junjie
Lee, Kong Aik
Truong, Duc-Tuan
Liu, Tianchi
Mak, Man-Wai
Sound
Audio and Speech Processing
There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Since individual speech frames do not contribute equally to the utterance-level representation, it is essential to estimate the importance or reliability of each frame. The xi-vector model addresses this by assigning different weights to frames based on uncertainty estimation. However, its uncertainty estimation model is implicitly trained through classification loss alone and does not consider the temporal relationships between frames, which may lead to suboptimal supervision. In this paper, we propose an improved architecture, xi+. Compared to xi-vector, xi+ incorporates a temporal attention module to capture frame-level uncertainty in a context-aware manner. In addition, we introduce a novel loss function, Stochastic Variance Loss, which explicitly supervises the learning of uncertainty. Results demonstrate consistent performance improvements of about 10\% on the VoxCeleb1-O set and 11\% on the NIST SRE 2024 evaluation set.
title Xi+: Uncertainty Supervision for Robust Speaker Embedding
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.05993