Saved in:
Bibliographic Details
Main Authors: Huang, Weiming, Huang, Qinghua, Ma, Liyan, Wang, Chuan
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2310.14016
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911804005810176
author Huang, Weiming
Huang, Qinghua
Ma, Liyan
Wang, Chuan
author_facet Huang, Weiming
Huang, Qinghua
Ma, Liyan
Wang, Chuan
contents Sound event localization and detection (SELD) involves sound event detection (SED) and direction of arrival (DoA) estimation tasks. SED mainly relies on temporal dependencies to distinguish different sound classes, while DoA estimation depends on spatial correlations to estimate source directions. This paper addresses the need to simultaneously extract spatial-temporal information in audio signals to improve SELD performance. A novel block, the sliding-window graph-former (SwG-former), is designed to learn temporal context information of sound events based on their spatial correlations. The SwG-former block transforms audio signals into a graph representation and constructs graph vertices to capture higher abstraction levels for spatial correlations. It uses different-sized sliding windows to adapt various sound event durations and aggregates temporal features with similar spatial information while incorporating multi-head self-attention (MHSA) to model global information. Furthermore, as the cornerstone of message passing, a robust Conv2dAgg function is proposed and embedded into the block to aggregate the features of neighbor vertices. As a result, a SwG-former model, which stacks the SwG-former blocks, demonstrates superior performance compared to recent advanced SELD models. The SwG-former block is also integrated into the event-independent network version 2 (EINV2), called SwG-EINV2, which surpasses the state-of-the-art (SOTA) methods under the same acoustic environment.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14016
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SwG-former: A Sliding-Window Graph Convolutional Network for Simultaneous Spatial-Temporal Information Extraction in Sound Event Localization and Detection
Huang, Weiming
Huang, Qinghua
Ma, Liyan
Wang, Chuan
Audio and Speech Processing
Sound event localization and detection (SELD) involves sound event detection (SED) and direction of arrival (DoA) estimation tasks. SED mainly relies on temporal dependencies to distinguish different sound classes, while DoA estimation depends on spatial correlations to estimate source directions. This paper addresses the need to simultaneously extract spatial-temporal information in audio signals to improve SELD performance. A novel block, the sliding-window graph-former (SwG-former), is designed to learn temporal context information of sound events based on their spatial correlations. The SwG-former block transforms audio signals into a graph representation and constructs graph vertices to capture higher abstraction levels for spatial correlations. It uses different-sized sliding windows to adapt various sound event durations and aggregates temporal features with similar spatial information while incorporating multi-head self-attention (MHSA) to model global information. Furthermore, as the cornerstone of message passing, a robust Conv2dAgg function is proposed and embedded into the block to aggregate the features of neighbor vertices. As a result, a SwG-former model, which stacks the SwG-former blocks, demonstrates superior performance compared to recent advanced SELD models. The SwG-former block is also integrated into the event-independent network version 2 (EINV2), called SwG-EINV2, which surpasses the state-of-the-art (SOTA) methods under the same acoustic environment.
title SwG-former: A Sliding-Window Graph Convolutional Network for Simultaneous Spatial-Temporal Information Extraction in Sound Event Localization and Detection
topic Audio and Speech Processing
url https://arxiv.org/abs/2310.14016