Exploring Text-Queried Sound Event Detection with Audio Source Separation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929669427691520 |
|---|---|
| author | Yin, Han Bai, Jisheng Xiao, Yang Wang, Hui Zheng, Siqi Chen, Yafeng Das, Rohan Kumar Deng, Chong Chen, Jianfeng |
| author_facet | Yin, Han Bai, Jisheng Xiao, Yang Wang, Hui Zheng, Siqi Chen, Yafeng Das, Rohan Kumar Deng, Chong Chen, Jianfeng |
| contents | In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22\% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_13292 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Exploring Text-Queried Sound Event Detection with Audio Source Separation Yin, Han Bai, Jisheng Xiao, Yang Wang, Hui Zheng, Siqi Chen, Yafeng Das, Rohan Kumar Deng, Chong Chen, Jianfeng Audio and Speech Processing Sound In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22\% on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https://github.com/apple-yinhan/TQ-SED. |
| title | Exploring Text-Queried Sound Event Detection with Audio Source Separation |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2409.13292 |