Mitigating Non-Target Speaker Bias in Guided Speaker Embedding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Horiguchi, Shota, Ashihara, Takanori, Delcroix, Marc, Ando, Atsushi, Tawara, Naohiro
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915343629287424
author Horiguchi, Shota
Ashihara, Takanori
Delcroix, Marc
Ando, Atsushi
Tawara, Naohiro
author_facet Horiguchi, Shota
Ashihara, Takanori
Delcroix, Marc
Ando, Atsushi
Tawara, Naohiro
contents Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
Horiguchi, Shota
Ashihara, Takanori
Delcroix, Marc
Ando, Atsushi
Tawara, Naohiro
Audio and Speech Processing
Sound
Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets.
title Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.12500