Sound Event Detection with Boundary-Aware Optimization and Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915714265251840 |
|---|---|
| author | Schmid, Florian Tang, Chi Ian Parekh, Sanjeel Ithapu, Vamsi Krishna Ortiz, Juan Azcarreta Ferroni, Giacomo Qian, Yijun Jasonas, Arnoldas Frateanu, Cosmin Clark, Camilla Widmer, Gerhard Bilen, Çağdaş |
| author_facet | Schmid, Florian Tang, Chi Ian Parekh, Sanjeel Ithapu, Vamsi Krishna Ortiz, Juan Azcarreta Ferroni, Giacomo Qian, Yijun Jasonas, Arnoldas Frateanu, Cosmin Clark, Camilla Widmer, Gerhard Bilen, Çağdaş |
| contents | Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers - Recurrent Event Detection (RED) and Event Proposal Network (EPN) - which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_04178 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Sound Event Detection with Boundary-Aware Optimization and Inference Schmid, Florian Tang, Chi Ian Parekh, Sanjeel Ithapu, Vamsi Krishna Ortiz, Juan Azcarreta Ferroni, Giacomo Qian, Yijun Jasonas, Arnoldas Frateanu, Cosmin Clark, Camilla Widmer, Gerhard Bilen, Çağdaş Audio and Speech Processing Sound Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers - Recurrent Event Detection (RED) and Event Proposal Network (EPN) - which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes. |
| title | Sound Event Detection with Boundary-Aware Optimization and Inference |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2601.04178 |