On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912770796027904 |
|---|---|
| author | Baroudi, Séverin Bredin, Hervé Razik, Joseph Marxer, Ricard |
| author_facet | Baroudi, Séverin Bredin, Hervé Razik, Joseph Marxer, Ricard |
| contents | Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource settings, over the past few years. Despite this, evaluations on tasks such as Speaker Diarization and Speech Separation remain limited. This paper investigates the quality of recent self-supervised speech representations on these two speaker identity-related tasks, highlighting gaps in the current literature that stem from limitations in the existing benchmarks, particularly the lack of diversity in evaluation datasets and variety in downstream systems associated to both diarization and separation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_15224 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation Baroudi, Séverin Bredin, Hervé Razik, Joseph Marxer, Ricard Audio and Speech Processing Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource settings, over the past few years. Despite this, evaluations on tasks such as Speaker Diarization and Speech Separation remain limited. This paper investigates the quality of recent self-supervised speech representations on these two speaker identity-related tasks, highlighting gaps in the current literature that stem from limitations in the existing benchmarks, particularly the lack of diversity in evaluation datasets and variety in downstream systems associated to both diarization and separation. |
| title | On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.15224 |