Multi-layer attentive probing improves transfer of audio representations for bioacoustics
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914553825067008 |
|---|---|
| author | Miron, Marius Robinson, David Hagiwara, Masato Parcollet, Titouan Cauzinille, Jules Narula, Gagan Alizadeh, Milad Gilsenan-McMahon, Ellen Keen, Sara Chemla, Emmanuel Hoffman, Benjamin Cusimano, Maddie Kim, Diane Effenberger, Felix Lawton, Jane K. Raskin, Aza Pietquin, Olivier Geist, Matthieu |
| author_facet | Miron, Marius Robinson, David Hagiwara, Masato Parcollet, Titouan Cauzinille, Jules Narula, Gagan Alizadeh, Milad Gilsenan-McMahon, Ellen Keen, Sara Chemla, Emmanuel Hoffman, Benjamin Cusimano, Maddie Kim, Diane Effenberger, Felix Lawton, Jane K. Raskin, Aza Pietquin, Olivier Geist, Matthieu |
| contents | Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-capacity probe, such as a linear layer on the final encoder layer. While this standardization enables model comparisons, it may bias results by overlooking the interaction between encoder features and probe design. In this work, we systematically study different probing strategies across two bioacoustic benchmarks, BEANs and BirdSet. We evaluate last- and multi-layer probing, across linear and attention probes. We show that larger probe heads that leverage time information have superior performance. Our results suggest that current benchmarks may misrepresent encoder quality when relying on a last-layer probing setup. Multi-layer probing improves downstream task performance across all tested models, while attention probing has superior performance to linear probing for transformer models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_10494 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Multi-layer attentive probing improves transfer of audio representations for bioacoustics Miron, Marius Robinson, David Hagiwara, Masato Parcollet, Titouan Cauzinille, Jules Narula, Gagan Alizadeh, Milad Gilsenan-McMahon, Ellen Keen, Sara Chemla, Emmanuel Hoffman, Benjamin Cusimano, Maddie Kim, Diane Effenberger, Felix Lawton, Jane K. Raskin, Aza Pietquin, Olivier Geist, Matthieu Sound Artificial Intelligence Probing heads map the representations learned from audio by a machine learning model to downstream task labels and are a key component in evaluating representation learning. Most bioacoustic benchmarks use a fixed, low-capacity probe, such as a linear layer on the final encoder layer. While this standardization enables model comparisons, it may bias results by overlooking the interaction between encoder features and probe design. In this work, we systematically study different probing strategies across two bioacoustic benchmarks, BEANs and BirdSet. We evaluate last- and multi-layer probing, across linear and attention probes. We show that larger probe heads that leverage time information have superior performance. Our results suggest that current benchmarks may misrepresent encoder quality when relying on a last-layer probing setup. Multi-layer probing improves downstream task performance across all tested models, while attention probing has superior performance to linear probing for transformer models. |
| title | Multi-layer attentive probing improves transfer of audio representations for bioacoustics |
| topic | Sound Artificial Intelligence |
| url | https://arxiv.org/abs/2605.10494 |