Sound Event Detection with Boundary-Aware Optimization and Inference

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Schmid, Florian, Tang, Chi Ian, Parekh, Sanjeel, Ithapu, Vamsi Krishna, Ortiz, Juan Azcarreta, Ferroni, Giacomo, Qian, Yijun, Jasonas, Arnoldas, Frateanu, Cosmin, Clark, Camilla, Widmer, Gerhard, Bilen, Çağdaş
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915714265251840
author Schmid, Florian
Tang, Chi Ian
Parekh, Sanjeel
Ithapu, Vamsi Krishna
Ortiz, Juan Azcarreta
Ferroni, Giacomo
Qian, Yijun
Jasonas, Arnoldas
Frateanu, Cosmin
Clark, Camilla
Widmer, Gerhard
Bilen, Çağdaş
author_facet Schmid, Florian
Tang, Chi Ian
Parekh, Sanjeel
Ithapu, Vamsi Krishna
Ortiz, Juan Azcarreta
Ferroni, Giacomo
Qian, Yijun
Jasonas, Arnoldas
Frateanu, Cosmin
Clark, Camilla
Widmer, Gerhard
Bilen, Çağdaş
contents Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers - Recurrent Event Detection (RED) and Event Proposal Network (EPN) - which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04178
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Sound Event Detection with Boundary-Aware Optimization and Inference
Schmid, Florian
Tang, Chi Ian
Parekh, Sanjeel
Ithapu, Vamsi Krishna
Ortiz, Juan Azcarreta
Ferroni, Giacomo
Qian, Yijun
Jasonas, Arnoldas
Frateanu, Cosmin
Clark, Camilla
Widmer, Gerhard
Bilen, Çağdaş
Audio and Speech Processing
Sound
Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers - Recurrent Event Detection (RED) and Event Proposal Network (EPN) - which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes.
title Sound Event Detection with Boundary-Aware Optimization and Inference
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2601.04178