Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Luyao, Zheng, Siqi, Zhang, Qinglin, Wang, Hui, Chen, Yafeng, Chen, Qian, Zhang, Shiliang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909092016029696
author Cheng, Luyao
Zheng, Siqi
Zhang, Qinglin
Wang, Hui
Chen, Yafeng
Chen, Qian
Zhang, Shiliang
author_facet Cheng, Luyao
Zheng, Siqi
Zhang, Qinglin
Wang, Hui
Chen, Yafeng
Chen, Qian
Zhang, Shiliang
contents Speaker diarization has gained considerable attention within speech processing research community. Mainstream speaker diarization rely primarily on speakers' voice characteristics extracted from acoustic signals and often overlook the potential of semantic information. Considering the fact that speech signals can efficiently convey the content of a speech, it is of our interest to fully exploit these semantic cues utilizing language models. In this work we propose a novel approach to effectively leverage semantic information in clustering-based speaker diarization systems. Firstly, we introduce spoken language understanding modules to extract speaker-related semantic information and utilize these information to construct pairwise constraints. Secondly, we present a novel framework to integrate these constraints into the speaker diarization pipeline, enhancing the performance of the entire system. Extensive experiments conducted on the public dataset demonstrate the consistent superiority of our proposed approach over acoustic-only speaker diarization systems.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10456
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
Cheng, Luyao
Zheng, Siqi
Zhang, Qinglin
Wang, Hui
Chen, Yafeng
Chen, Qian
Zhang, Shiliang
Sound
Computation and Language
Audio and Speech Processing
Speaker diarization has gained considerable attention within speech processing research community. Mainstream speaker diarization rely primarily on speakers' voice characteristics extracted from acoustic signals and often overlook the potential of semantic information. Considering the fact that speech signals can efficiently convey the content of a speech, it is of our interest to fully exploit these semantic cues utilizing language models. In this work we propose a novel approach to effectively leverage semantic information in clustering-based speaker diarization systems. Firstly, we introduce spoken language understanding modules to extract speaker-related semantic information and utilize these information to construct pairwise constraints. Secondly, we present a novel framework to integrate these constraints into the speaker diarization pipeline, enhancing the performance of the entire system. Extensive experiments conducted on the public dataset demonstrate the consistent superiority of our proposed approach over acoustic-only speaker diarization systems.
title Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2309.10456