Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908664667832320 |
|---|---|
| author | Marincione, Davide Crisostomi, Donato Dessi, Roberto Rodolà, Emanuele Rossi, Emanuele |
| author_facet | Marincione, Davide Crisostomi, Donato Dessi, Roberto Rodolà, Emanuele Rossi, Emanuele |
| contents | Foundation models capable of generalizing across species and tasks represent a promising new frontier in bioacoustics, with NatureLM being one of the most prominent examples. While its domain-specific fine-tuning yields strong performance on bioacoustic benchmarks, we observe that it also introduces trade-offs in instruction-following flexibility. For instance, NatureLM achieves high accuracy when prompted for either the common or scientific name individually, but its accuracy drops significantly when both are requested in a single prompt. We address this by applying a simple model merging strategy that interpolates NatureLM with its base language model, recovering instruction-following capabilities with minimal loss of domain expertise. Finally, we show that the merged model exhibits markedly stronger zero-shot generalization, achieving over a 200% relative improvement and setting a new state-of-the-art in closed-set zero-shot classification of unseen species. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_05171 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models Marincione, Davide Crisostomi, Donato Dessi, Roberto Rodolà, Emanuele Rossi, Emanuele Machine Learning Artificial Intelligence Sound Foundation models capable of generalizing across species and tasks represent a promising new frontier in bioacoustics, with NatureLM being one of the most prominent examples. While its domain-specific fine-tuning yields strong performance on bioacoustic benchmarks, we observe that it also introduces trade-offs in instruction-following flexibility. For instance, NatureLM achieves high accuracy when prompted for either the common or scientific name individually, but its accuracy drops significantly when both are requested in a single prompt. We address this by applying a simple model merging strategy that interpolates NatureLM with its base language model, recovering instruction-following capabilities with minimal loss of domain expertise. Finally, we show that the merged model exhibits markedly stronger zero-shot generalization, achieving over a 200% relative improvement and setting a new state-of-the-art in closed-set zero-shot classification of unseen species. |
| title | Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models |
| topic | Machine Learning Artificial Intelligence Sound |
| url | https://arxiv.org/abs/2511.05171 |