Robust detection of overlapping bioacoustic sound events

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mahon, Louis, Hoffman, Benjamin, James, Logan, Cusimano, Maddie, Hagiwara, Masato, Woolley, Sarah C, Effenberger, Felix, Keen, Sara, Liu, Jen-Yu, Pietquin, Olivier
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916936547303424
author Mahon, Louis
Hoffman, Benjamin
James, Logan
Cusimano, Maddie
Hagiwara, Masato
Woolley, Sarah C
Effenberger, Felix
Keen, Sara
Liu, Jen-Yu
Pietquin, Olivier
author_facet Mahon, Louis
Hoffman, Benjamin
James, Logan
Cusimano, Maddie
Hagiwara, Masato
Woolley, Sarah C
Effenberger, Felix
Keen, Sara
Liu, Jen-Yu
Pietquin, Olivier
contents We propose a method for accurately detecting bioacoustic sound events that is robust to overlapping events, a common issue in domains such as ethology, ecology and conservation. While standard methods employ a frame-based, multi-label approach, we introduce an onset-based detection method which we name Voxaboxen. It takes inspiration from object detection methods in computer vision, but simultaneously takes advantage of recent advances in self-supervised audio encoders. For each time window, Voxaboxen predicts whether it contains the start of a vocalization and how long the vocalization is. It also does the same in reverse, predicting whether each window contains the end of a vocalization, and how long ago it started. The two resulting sets of bounding boxes are then fused using a graph-matching algorithm. We also release a new dataset designed to measure performance on detecting overlapping vocalizations. This consists of recordings of zebra finches annotated with temporally-strong labels and showing frequent overlaps. We test Voxaboxen on seven existing data sets and on our new data set. We compare Voxaboxen to natural baselines and existing sound event detection methods and demonstrate SotA results. Further experiments show that improvements are robust to frequent vocalization overlap.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust detection of overlapping bioacoustic sound events
Mahon, Louis
Hoffman, Benjamin
James, Logan
Cusimano, Maddie
Hagiwara, Masato
Woolley, Sarah C
Effenberger, Felix
Keen, Sara
Liu, Jen-Yu
Pietquin, Olivier
Sound
Machine Learning
Audio and Speech Processing
We propose a method for accurately detecting bioacoustic sound events that is robust to overlapping events, a common issue in domains such as ethology, ecology and conservation. While standard methods employ a frame-based, multi-label approach, we introduce an onset-based detection method which we name Voxaboxen. It takes inspiration from object detection methods in computer vision, but simultaneously takes advantage of recent advances in self-supervised audio encoders. For each time window, Voxaboxen predicts whether it contains the start of a vocalization and how long the vocalization is. It also does the same in reverse, predicting whether each window contains the end of a vocalization, and how long ago it started. The two resulting sets of bounding boxes are then fused using a graph-matching algorithm. We also release a new dataset designed to measure performance on detecting overlapping vocalizations. This consists of recordings of zebra finches annotated with temporally-strong labels and showing frequent overlaps. We test Voxaboxen on seven existing data sets and on our new data set. We compare Voxaboxen to natural baselines and existing sound event detection methods and demonstrate SotA results. Further experiments show that improvements are robust to frequent vocalization overlap.
title Robust detection of overlapping bioacoustic sound events
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2503.02389