AVEX: What Matters for Animal Vocalization Encoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miron, Marius, Robinson, David, Alizadeh, Milad, Gilsenan-McMahon, Ellen, Narula, Gagan, Chemla, Emmanuel, Cusimano, Maddie, Effenberger, Felix, Hagiwara, Masato, Hoffman, Benjamin, Keen, Sara, Kim, Diane, Lawton, Jane, Liu, Jen-Yu, Raskin, Aza, Pietquin, Olivier, Geist, Matthieu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909042631245824
author Miron, Marius
Robinson, David
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Narula, Gagan
Chemla, Emmanuel
Cusimano, Maddie
Effenberger, Felix
Hagiwara, Masato
Hoffman, Benjamin
Keen, Sara
Kim, Diane
Lawton, Jane
Liu, Jen-Yu
Raskin, Aza
Pietquin, Olivier
Geist, Matthieu
author_facet Miron, Marius
Robinson, David
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Narula, Gagan
Chemla, Emmanuel
Cusimano, Maddie
Effenberger, Felix
Hagiwara, Masato
Hoffman, Benjamin
Keen, Sara
Kim, Diane
Lawton, Jane
Liu, Jen-Yu
Raskin, Aza
Pietquin, Olivier
Geist, Matthieu
contents Bioacoustics, the study of sounds produced by living organisms, plays a vital role in conservation, biodiversity monitoring, and behavioral studies. Many tasks in this field, such as species, individual, and behavior classification and detection, are well-suited to machine learning. However, they often suffer from limited annotated data, highlighting the need for a general-purpose bioacoustic encoder capable of extracting useful representations for diverse downstream tasks. Such encoders have been proposed before, but are often limited in scope due to a focus on a narrow range of species (typically birds), and a reliance on a single model architecture or training paradigm. Moreover, they are usually evaluated on a small set of tasks and datasets. In this work, we present a large-scale empirical study that covers aspects of bioacoustics that are relevant to research but have previously been scarcely considered: training data diversity and scale, model architectures and training recipes, and the breadth of evaluation tasks and datasets. We obtain encoders that are state-of-the-art on the existing and proposed benchmarks. We also identify what matters for training these encoders, such that this work can be extended when more data are available or better architectures are proposed. Specifically, across 26 datasets with tasks including species classification, detection, individual ID, and vocal repertoire discovery, we find self-supervised pre-training followed by supervised post-training on a mixed bioacoustics + general-audio corpus yields the strongest in- and out-of-distribution performance. We show the importance of data diversity in both stages. To support ongoing research and application, we will release the model checkpoints.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AVEX: What Matters for Animal Vocalization Encoding
Miron, Marius
Robinson, David
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Narula, Gagan
Chemla, Emmanuel
Cusimano, Maddie
Effenberger, Felix
Hagiwara, Masato
Hoffman, Benjamin
Keen, Sara
Kim, Diane
Lawton, Jane
Liu, Jen-Yu
Raskin, Aza
Pietquin, Olivier
Geist, Matthieu
Sound
Artificial Intelligence
Information Retrieval
Machine Learning
Bioacoustics, the study of sounds produced by living organisms, plays a vital role in conservation, biodiversity monitoring, and behavioral studies. Many tasks in this field, such as species, individual, and behavior classification and detection, are well-suited to machine learning. However, they often suffer from limited annotated data, highlighting the need for a general-purpose bioacoustic encoder capable of extracting useful representations for diverse downstream tasks. Such encoders have been proposed before, but are often limited in scope due to a focus on a narrow range of species (typically birds), and a reliance on a single model architecture or training paradigm. Moreover, they are usually evaluated on a small set of tasks and datasets. In this work, we present a large-scale empirical study that covers aspects of bioacoustics that are relevant to research but have previously been scarcely considered: training data diversity and scale, model architectures and training recipes, and the breadth of evaluation tasks and datasets. We obtain encoders that are state-of-the-art on the existing and proposed benchmarks. We also identify what matters for training these encoders, such that this work can be extended when more data are available or better architectures are proposed. Specifically, across 26 datasets with tasks including species classification, detection, individual ID, and vocal repertoire discovery, we find self-supervised pre-training followed by supervised post-training on a mixed bioacoustics + general-audio corpus yields the strongest in- and out-of-distribution performance. We show the importance of data diversity in both stages. To support ongoing research and application, we will release the model checkpoints.
title AVEX: What Matters for Animal Vocalization Encoding
topic Sound
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2508.11845