Text-Queried Target Sound Event Localization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913402437238784 |
|---|---|
| author | Zhao, Jinzheng Qian, Xinyuan Xu, Yong Liu, Haohe Cao, Yin Berghi, Davide Wang, Wenwu |
| author_facet | Zhao, Jinzheng Qian, Xinyuan Xu, Yong Liu, Haohe Cao, Yin Berghi, Davide Wang, Wenwu |
| contents | Sound event localization and detection (SELD) aims to determine the appearance of sound classes, together with their Direction of Arrival (DOA). However, current SELD systems can only predict the activities of specific classes, for example, 13 classes in DCASE challenges. In this paper, we propose text-queried target sound event localization (SEL), a new paradigm that allows the user to input the text to describe the sound event, and the SEL model can predict the location of the related sound event. The proposed task presents a more user-friendly way for human-computer interaction. We provide a benchmark study for the proposed task and perform experiments on datasets created by simulated room impulse response (RIR) and real RIR to validate the effectiveness of the proposed methods. We hope that our benchmark will inspire the interest and additional research for text-queried sound source localization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_16058 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Text-Queried Target Sound Event Localization Zhao, Jinzheng Qian, Xinyuan Xu, Yong Liu, Haohe Cao, Yin Berghi, Davide Wang, Wenwu Audio and Speech Processing Sound event localization and detection (SELD) aims to determine the appearance of sound classes, together with their Direction of Arrival (DOA). However, current SELD systems can only predict the activities of specific classes, for example, 13 classes in DCASE challenges. In this paper, we propose text-queried target sound event localization (SEL), a new paradigm that allows the user to input the text to describe the sound event, and the SEL model can predict the location of the related sound event. The proposed task presents a more user-friendly way for human-computer interaction. We provide a benchmark study for the proposed task and perform experiments on datasets created by simulated room impulse response (RIR) and real RIR to validate the effectiveness of the proposed methods. We hope that our benchmark will inspire the interest and additional research for text-queried sound source localization. |
| title | Text-Queried Target Sound Event Localization |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.16058 |