Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Iatariene, Taous, Guérin, Alexandre, Serizel, Romain
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912545010352128
author Iatariene, Taous
Guérin, Alexandre
Serizel, Romain
author_facet Iatariene, Taous
Guérin, Alexandre
Serizel, Romain
contents Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding extractors usually struggle with short temporal contexts and overlapping speech, which imposes long-term identity reassignment to exploit longer temporal contexts. However, this increases the probability of tracking system errors, which in turn impacts negatively on identity reassignment. To address this, we propose a Knowledge Distillation (KD) based training approach for short context speaker embedding extraction from two speaker mixtures. We leverage the spatial information of the speaker of interest using beamforming to reduce overlap. We study the feasibility of performing identity reassignment over blocks of fixed size, i.e., blockwise identity reassignment, to go towards a low-latency speaker embedding based tracking system. Results demonstrate that our distilled models are effective at short-context embedding extraction and more robust to overlap. Although, blockwise reassignment results indicate that further work is needed to handle simultaneous speech more effectively.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
Iatariene, Taous
Guérin, Alexandre
Serizel, Romain
Audio and Speech Processing
Artificial Intelligence
Sound
Signal Processing
Speaker embeddings are promising identity-related features that can enhance the identity assignment performance of a tracking system by leveraging its spatial predictions, i.e, by performing identity reassignment. Common speaker embedding extractors usually struggle with short temporal contexts and overlapping speech, which imposes long-term identity reassignment to exploit longer temporal contexts. However, this increases the probability of tracking system errors, which in turn impacts negatively on identity reassignment. To address this, we propose a Knowledge Distillation (KD) based training approach for short context speaker embedding extraction from two speaker mixtures. We leverage the spatial information of the speaker of interest using beamforming to reduce overlap. We study the feasibility of performing identity reassignment over blocks of fixed size, i.e., blockwise identity reassignment, to go towards a low-latency speaker embedding based tracking system. Results demonstrate that our distilled models are effective at short-context embedding extraction and more robust to overlap. Although, blockwise reassignment results indicate that further work is needed to handle simultaneous speech more effectively.
title Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
topic Audio and Speech Processing
Artificial Intelligence
Sound
Signal Processing
url https://arxiv.org/abs/2508.14115