Text-Queried Target Sound Event Localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jinzheng, Qian, Xinyuan, Xu, Yong, Liu, Haohe, Cao, Yin, Berghi, Davide, Wang, Wenwu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913402437238784
author Zhao, Jinzheng
Qian, Xinyuan
Xu, Yong
Liu, Haohe
Cao, Yin
Berghi, Davide
Wang, Wenwu
author_facet Zhao, Jinzheng
Qian, Xinyuan
Xu, Yong
Liu, Haohe
Cao, Yin
Berghi, Davide
Wang, Wenwu
contents Sound event localization and detection (SELD) aims to determine the appearance of sound classes, together with their Direction of Arrival (DOA). However, current SELD systems can only predict the activities of specific classes, for example, 13 classes in DCASE challenges. In this paper, we propose text-queried target sound event localization (SEL), a new paradigm that allows the user to input the text to describe the sound event, and the SEL model can predict the location of the related sound event. The proposed task presents a more user-friendly way for human-computer interaction. We provide a benchmark study for the proposed task and perform experiments on datasets created by simulated room impulse response (RIR) and real RIR to validate the effectiveness of the proposed methods. We hope that our benchmark will inspire the interest and additional research for text-queried sound source localization.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16058
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text-Queried Target Sound Event Localization
Zhao, Jinzheng
Qian, Xinyuan
Xu, Yong
Liu, Haohe
Cao, Yin
Berghi, Davide
Wang, Wenwu
Audio and Speech Processing
Sound event localization and detection (SELD) aims to determine the appearance of sound classes, together with their Direction of Arrival (DOA). However, current SELD systems can only predict the activities of specific classes, for example, 13 classes in DCASE challenges. In this paper, we propose text-queried target sound event localization (SEL), a new paradigm that allows the user to input the text to describe the sound event, and the SEL model can predict the location of the related sound event. The proposed task presents a more user-friendly way for human-computer interaction. We provide a benchmark study for the proposed task and perform experiments on datasets created by simulated room impulse response (RIR) and real RIR to validate the effectiveness of the proposed methods. We hope that our benchmark will inspire the interest and additional research for text-queried sound source localization.
title Text-Queried Target Sound Event Localization
topic Audio and Speech Processing
url https://arxiv.org/abs/2406.16058