Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Zhaokai, Zhang, Li, Wang, Qing, Zhou, Pan, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912402024431616
author Sun, Zhaokai
Zhang, Li
Wang, Qing
Zhou, Pan
Xie, Lei
author_facet Sun, Zhaokai
Zhang, Li
Wang, Qing
Zhou, Pan
Xie, Lei
contents Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76\% on the AMI test set, demonstrating its robustness and effectiveness in OSD.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
Sun, Zhaokai
Zhang, Li
Wang, Qing
Zhou, Pan
Xie, Lei
Sound
Machine Learning
Audio and Speech Processing
Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76\% on the AMI test set, demonstrating its robustness and effectiveness in OSD.
title Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2505.23207