Exploring Text-Queried Sound Event Detection with Audio Source Separation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yin, Han, Bai, Jisheng, Xiao, Yang, Wang, Hui, Zheng, Siqi, Chen, Yafeng, Das, Rohan Kumar, Deng, Chong, Chen, Jianfeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929669427691520
author Yin, Han
Bai, Jisheng
Xiao, Yang
Wang, Hui
Zheng, Siqi
Chen, Yafeng
Das, Rohan Kumar
Deng, Chong
Chen, Jianfeng
author_facet Yin, Han
Bai, Jisheng
Xiao, Yang
Wang, Hui
Zheng, Siqi
Chen, Yafeng
Das, Rohan Kumar
Deng, Chong
Chen, Jianfeng
contents In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22\% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Text-Queried Sound Event Detection with Audio Source Separation
Yin, Han
Bai, Jisheng
Xiao, Yang
Wang, Hui
Zheng, Siqi
Chen, Yafeng
Das, Rohan Kumar
Deng, Chong
Chen, Jianfeng
Audio and Speech Processing
Sound
In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22\% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED.
title Exploring Text-Queried Sound Event Detection with Audio Source Separation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.13292