Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915343629287424 |
|---|---|
| author | Horiguchi, Shota Ashihara, Takanori Delcroix, Marc Ando, Atsushi Tawara, Naohiro |
| author_facet | Horiguchi, Shota Ashihara, Takanori Delcroix, Marc Ando, Atsushi Tawara, Naohiro |
| contents | Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_12500 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Mitigating Non-Target Speaker Bias in Guided Speaker Embedding Horiguchi, Shota Ashihara, Takanori Delcroix, Marc Ando, Atsushi Tawara, Naohiro Audio and Speech Processing Sound Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues, drastically improved embeddings under severe overlap with small degradation in low-overlap cases. However, since extreme overlaps are rare in natural conversations, this degradation cannot be overlooked. This paper first reveals that the degradation is caused by the global-statistics-based modules, widely used in speaker embedding extractors, being overly sensitive to intervals containing only non-target speakers. As a countermeasure, we propose an extension of such modules that exploit the target speaker activity clues, to compute statistics from intervals where the target is active. The proposed method improves speaker verification performance in both low and high overlap ratios, and diarization performance on multiple datasets. |
| title | Mitigating Non-Target Speaker Bias in Guided Speaker Embedding |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2506.12500 |