From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Mengcheng, Zhou, Xue, Xu, Chen, Man, Dapeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911398966067200
author Huang, Mengcheng
Zhou, Xue
Xu, Chen
Man, Dapeng
author_facet Huang, Mengcheng
Zhou, Xue
Xu, Chen
Man, Dapeng
contents Underwater acoustic target recognition (UATR) plays a vital role in marine applications but remains challenging due to limited labeled data and the complexity of ocean environments. This paper explores a central question: can speech large models (SLMs), trained on massive human speech corpora, be effectively transferred to underwater acoustics? To investigate this, we propose UATR-SLM, a simple framework that reuses the speech feature pipeline, adapts the SLM as an acoustic encoder, and adds a lightweight classifier.Experiments on the DeepShip and ShipsEar benchmarks show that UATR-SLM achieves over 99% in-domain accuracy, maintains strong robustness across variable signal lengths, and reaches up to 96.67% accuracy in cross-domain evaluation. These results highlight the strong transferability of SLMs to UATR, establishing a promising paradigm for leveraging speech foundation models in underwater acoustics.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18086
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
Huang, Mengcheng
Zhou, Xue
Xu, Chen
Man, Dapeng
Sound
Audio and Speech Processing
Underwater acoustic target recognition (UATR) plays a vital role in marine applications but remains challenging due to limited labeled data and the complexity of ocean environments. This paper explores a central question: can speech large models (SLMs), trained on massive human speech corpora, be effectively transferred to underwater acoustics? To investigate this, we propose UATR-SLM, a simple framework that reuses the speech feature pipeline, adapts the SLM as an acoustic encoder, and adds a lightweight classifier.Experiments on the DeepShip and ShipsEar benchmarks show that UATR-SLM achieves over 99% in-domain accuracy, maintains strong robustness across variable signal lengths, and reaches up to 96.67% accuracy in cross-domain evaluation. These results highlight the strong transferability of SLMs to UATR, establishing a promising paradigm for leveraging speech foundation models in underwater acoustics.
title From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.18086