Aggregation Strategies for Efficient Annotation of Bioacoustic Sound Events Using Active Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lindholm, Richard, Marklund, Oscar, Mogren, Olof, Martinsson, John
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912257895563264
author Lindholm, Richard
Marklund, Oscar
Mogren, Olof
Martinsson, John
author_facet Lindholm, Richard
Marklund, Oscar
Mogren, Olof
Martinsson, John
contents The vast amounts of audio data collected in Sound Event Detection (SED) applications require efficient annotation strategies to enable supervised learning. Manual labeling is expensive and time-consuming, making Active Learning (AL) a promising approach for reducing annotation effort. We introduce Top K Entropy, a novel uncertainty aggregation strategy for AL that prioritizes the most uncertain segments within an audio recording, instead of averaging uncertainty across all segments. This approach enables the selection of entire recordings for annotation, improving efficiency in sparse data scenarios. We compare Top K Entropy to random sampling and Mean Entropy, and show that fewer labels can lead to the same model performance, particularly in datasets with sparse sound events. Evaluations are conducted on audio mixtures of sound recordings from parks with meerkat, dog, and baby crying sound events, representing real-world bioacoustic monitoring scenarios. Using Top K Entropy for active learning, we can achieve comparable performance to training on the fully labeled dataset with only 8% of the labels. Top K Entropy outperforms Mean Entropy, suggesting that it is best to let the most uncertain segments represent the uncertainty of an audio file. The findings highlight the potential of AL for scalable annotation in audio and time-series applications, including bioacoustics.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02422
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aggregation Strategies for Efficient Annotation of Bioacoustic Sound Events Using Active Learning
Lindholm, Richard
Marklund, Oscar
Mogren, Olof
Martinsson, John
Sound
Machine Learning
Audio and Speech Processing
The vast amounts of audio data collected in Sound Event Detection (SED) applications require efficient annotation strategies to enable supervised learning. Manual labeling is expensive and time-consuming, making Active Learning (AL) a promising approach for reducing annotation effort. We introduce Top K Entropy, a novel uncertainty aggregation strategy for AL that prioritizes the most uncertain segments within an audio recording, instead of averaging uncertainty across all segments. This approach enables the selection of entire recordings for annotation, improving efficiency in sparse data scenarios. We compare Top K Entropy to random sampling and Mean Entropy, and show that fewer labels can lead to the same model performance, particularly in datasets with sparse sound events. Evaluations are conducted on audio mixtures of sound recordings from parks with meerkat, dog, and baby crying sound events, representing real-world bioacoustic monitoring scenarios. Using Top K Entropy for active learning, we can achieve comparable performance to training on the fully labeled dataset with only 8% of the labels. Top K Entropy outperforms Mean Entropy, suggesting that it is best to let the most uncertain segments represent the uncertainty of an audio file. The findings highlight the potential of AL for scalable annotation in audio and time-series applications, including bioacoustics.
title Aggregation Strategies for Efficient Annotation of Bioacoustic Sound Events Using Active Learning
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2503.02422