Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912402024431616 |
|---|---|
| author | Sun, Zhaokai Zhang, Li Wang, Qing Zhou, Pan Xie, Lei |
| author_facet | Sun, Zhaokai Zhang, Li Wang, Qing Zhou, Pan Xie, Lei |
| contents | Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76\% on the AMI test set, demonstrating its robustness and effectiveness in OSD. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_23207 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM Sun, Zhaokai Zhang, Li Wang, Qing Zhou, Pan Xie, Lei Sound Machine Learning Audio and Speech Processing Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76\% on the AMI test set, demonstrating its robustness and effectiveness in OSD. |
| title | Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.23207 |