Sound event detection with audio-text models and heterogeneous temporal annotations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Harju, Manu, Mesaros, Annamaria
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909757882761216
author Harju, Manu
Mesaros, Annamaria
author_facet Harju, Manu
Mesaros, Annamaria
contents Recent advances in generating synthetic captions based on audio and related metadata allow using the information contained in natural language as input for other audio tasks. In this paper, we propose a novel method to guide a sound event detection system with free-form text. We use machine-generated captions as complementary information to the strong labels for training, and evaluate the systems using different types of textual inputs. In addition, we study a scenario where only part of the training data has strong labels, and the rest of it only has temporally weak labels. Our findings show that synthetic captions improve the performance in both cases compared to the CRNN architecture typically used for sound event detection. On a dataset of 50 highly unbalanced classes, the PSDS-1 score increases from 0.223 to 0.277 when trained with strong labels, and from 0.166 to 0.218 when half of the training data has only weak labels.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20703
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sound event detection with audio-text models and heterogeneous temporal annotations
Harju, Manu
Mesaros, Annamaria
Audio and Speech Processing
Recent advances in generating synthetic captions based on audio and related metadata allow using the information contained in natural language as input for other audio tasks. In this paper, we propose a novel method to guide a sound event detection system with free-form text. We use machine-generated captions as complementary information to the strong labels for training, and evaluate the systems using different types of textual inputs. In addition, we study a scenario where only part of the training data has strong labels, and the rest of it only has temporally weak labels. Our findings show that synthetic captions improve the performance in both cases compared to the CRNN architecture typically used for sound event detection. On a dataset of 50 highly unbalanced classes, the PSDS-1 score increases from 0.223 to 0.277 when trained with strong labels, and from 0.166 to 0.218 when half of the training data has only weak labels.
title Sound event detection with audio-text models and heterogeneous temporal annotations
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.20703