Multi-layer attentive probing improves transfer of audio representations for bioacoustics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miron, Marius, Robinson, David, Hagiwara, Masato, Parcollet, Titouan, Cauzinille, Jules, Narula, Gagan, Alizadeh, Milad, Gilsenan-McMahon, Ellen, Keen, Sara, Chemla, Emmanuel, Hoffman, Benjamin, Cusimano, Maddie, Kim, Diane, Effenberger, Felix, Lawton, Jane K., Raskin, Aza, Pietquin, Olivier, Geist, Matthieu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914553825067008
author Miron, Marius
Robinson, David
Hagiwara, Masato
Parcollet, Titouan
Cauzinille, Jules
Narula, Gagan
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Keen, Sara
Chemla, Emmanuel
Hoffman, Benjamin
Cusimano, Maddie
Kim, Diane
Effenberger, Felix
Lawton, Jane K.
Raskin, Aza
Pietquin, Olivier
Geist, Matthieu
author_facet Miron, Marius
Robinson, David
Hagiwara, Masato
Parcollet, Titouan
Cauzinille, Jules
Narula, Gagan
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Keen, Sara
Chemla, Emmanuel
Hoffman, Benjamin
Cusimano, Maddie
Kim, Diane
Effenberger, Felix
Lawton, Jane K.
Raskin, Aza
Pietquin, Olivier
Geist, Matthieu
contents Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-capacity probe, such as a linear layer on the final encoder layer. While this standardization enables model comparisons, it may bias results by overlooking the interaction between encoder features and probe design. In this work, we systematically study different probing strategies across two bioacoustic benchmarks, BEANs and BirdSet. We evaluate last- and multi-layer probing, across linear and attention probes. We show that larger probe heads that leverage time information have superior performance. Our results suggest that current benchmarks may misrepresent encoder quality when relying on a last-layer probing setup. Multi-layer probing improves downstream task performance across all tested models, while attention probing has superior performance to linear probing for transformer models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10494
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-layer attentive probing improves transfer of audio representations for bioacoustics
Miron, Marius
Robinson, David
Hagiwara, Masato
Parcollet, Titouan
Cauzinille, Jules
Narula, Gagan
Alizadeh, Milad
Gilsenan-McMahon, Ellen
Keen, Sara
Chemla, Emmanuel
Hoffman, Benjamin
Cusimano, Maddie
Kim, Diane
Effenberger, Felix
Lawton, Jane K.
Raskin, Aza
Pietquin, Olivier
Geist, Matthieu
Sound
Artificial Intelligence
Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-capacity probe, such as a linear layer on the final encoder layer. While this standardization enables model comparisons, it may bias results by overlooking the interaction between encoder features and probe design. In this work, we systematically study different probing strategies across two bioacoustic benchmarks, BEANs and BirdSet. We evaluate last- and multi-layer probing, across linear and attention probes. We show that larger probe heads that leverage time information have superior performance. Our results suggest that current benchmarks may misrepresent encoder quality when relying on a last-layer probing setup. Multi-layer probing improves downstream task performance across all tested models, while attention probing has superior performance to linear probing for transformer models.
title Multi-layer attentive probing improves transfer of audio representations for bioacoustics
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2605.10494