Target Speech Diarization with Multimodal Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yidi, Tao, Ruijie, Chen, Zhengyang, Qian, Yanmin, Li, Haizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911913632333824
author Jiang, Yidi
Tao, Ruijie
Chen, Zhengyang
Qian, Yanmin
Li, Haizhou
author_facet Jiang, Yidi
Tao, Ruijie
Chen, Zhengyang
Qian, Yanmin
Li, Haizhou
contents Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We propose a novel Multimodal Target Speech Diarization (MM-TSD) framework, which accommodates diverse and multi-modal prompts to specify target events in a flexible and user-friendly manner, including semantic language description, pre-enrolled speech, pre-registered face image, and audio-language logical prompts. We further propose a voice-face aligner module to project human voice and face representation into a shared space. We develop a multi-modal dataset based on VoxCeleb2 for MM-TSD training and evaluation. Additionally, we conduct comparative analysis and ablation studies for each category of prompts to validate the efficacy of each component in the proposed framework. Furthermore, our framework demonstrates versatility in performing various signal processing tasks, including speaker diarization and overlap speech detection, using task-specific prompts. MM-TSD achieves robust and comparable performance as a unified system compared to specialized models. Moreover, MM-TSD shows capability to handle complex conversations for real-world dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07198
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Target Speech Diarization with Multimodal Prompts
Jiang, Yidi
Tao, Ruijie
Chen, Zhengyang
Qian, Yanmin
Li, Haizhou
Audio and Speech Processing
Multimedia
Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We propose a novel Multimodal Target Speech Diarization (MM-TSD) framework, which accommodates diverse and multi-modal prompts to specify target events in a flexible and user-friendly manner, including semantic language description, pre-enrolled speech, pre-registered face image, and audio-language logical prompts. We further propose a voice-face aligner module to project human voice and face representation into a shared space. We develop a multi-modal dataset based on VoxCeleb2 for MM-TSD training and evaluation. Additionally, we conduct comparative analysis and ablation studies for each category of prompts to validate the efficacy of each component in the proposed framework. Furthermore, our framework demonstrates versatility in performing various signal processing tasks, including speaker diarization and overlap speech detection, using task-specific prompts. MM-TSD achieves robust and comparable performance as a unified system compared to specialized models. Moreover, MM-TSD shows capability to handle complex conversations for real-world dataset.
title Target Speech Diarization with Multimodal Prompts
topic Audio and Speech Processing
Multimedia
url https://arxiv.org/abs/2406.07198