Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Binh Thien, Yasuda, Masahiro, Takeuchi, Daiki, Niizumi, Daisuke, Ohishi, Yasunori, Harada, Noboru
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909642203856896
author Nguyen, Binh Thien
Yasuda, Masahiro
Takeuchi, Daiki
Niizumi, Daisuke
Ohishi, Yasunori
Harada, Noboru
author_facet Nguyen, Binh Thien
Yasuda, Masahiro
Takeuchi, Daiki
Niizumi, Daisuke
Ohishi, Yasunori
Harada, Noboru
contents Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for spatial semantic segmentation of sound scenes (S5), which focuses on detecting and separating sound events in spatial sound scenes. In this paper, we explore methods for addressing the S5 task. Specifically, we present baseline S5 systems that combine audio tagging (AT) and label-queried source separation (LSS) models. We investigate two LSS approaches based on the ResUNet architecture: a) extracting a single source for each detected event and b) querying multiple sources concurrently. Since each separated source in S5 is identified by its sound event class label, we propose new class-aware metrics to evaluate both the sound sources and labels simultaneously. Experimental results on first-order ambisonics spatial audio demonstrate the effectiveness of the proposed systems and confirm the efficacy of the metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
Nguyen, Binh Thien
Yasuda, Masahiro
Takeuchi, Daiki
Niizumi, Daisuke
Ohishi, Yasunori
Harada, Noboru
Audio and Speech Processing
Sound
Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for spatial semantic segmentation of sound scenes (S5), which focuses on detecting and separating sound events in spatial sound scenes. In this paper, we explore methods for addressing the S5 task. Specifically, we present baseline S5 systems that combine audio tagging (AT) and label-queried source separation (LSS) models. We investigate two LSS approaches based on the ResUNet architecture: a) extracting a single source for each detected event and b) querying multiple sources concurrently. Since each separated source in S5 is identified by its sound event class label, we propose new class-aware metrics to evaluate both the sound sources and labels simultaneously. Experimental results on first-order ambisonics spatial audio demonstrate the effectiveness of the proposed systems and confirm the efficacy of the metrics.
title Baseline Systems and Evaluation Metrics for Spatial Semantic Segmentation of Sound Scenes
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2503.22088