AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Yiwei, Li, Bohan, Wang, Hankun, Li, Zhihan, Wang, Shuai, Chen, Xie, Yu, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911293421649920
author Guo, Yiwei
Li, Bohan
Wang, Hankun
Li, Zhihan
Wang, Shuai
Chen, Xie
Yu, Kai
author_facet Guo, Yiwei
Li, Bohan
Wang, Hankun
Li, Zhihan
Wang, Shuai
Chen, Xie
Yu, Kai
contents Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we propose AHAMask, where we simply mask some of the attention heads in the decoder-only LLM backbone of LALMs, to trigger specific acoustic task functionalities without instructions. These masks are efficiently obtained by training on an LALM, with the number of trainable parameters equal to the attention head count in its LLM backbone. We show by experiments that applying such selective attention head masks achieves comparable or even better performance than using instructions, either on single or composite tasks. Besides achieving reliable acoustic task specification for LALMs, this also reveals that LALMs exhibit certain "functional pathways" in their attention heads.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01787
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
Guo, Yiwei
Li, Bohan
Wang, Hankun
Li, Zhihan
Wang, Shuai
Chen, Xie
Yu, Kai
Audio and Speech Processing
Artificial Intelligence
Sound
Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we propose AHAMask, where we simply mask some of the attention heads in the decoder-only LLM backbone of LALMs, to trigger specific acoustic task functionalities without instructions. These masks are efficiently obtained by training on an LALM, with the number of trainable parameters equal to the attention head count in its LLM backbone. We show by experiments that applying such selective attention head masks achieves comparable or even better performance than using instructions, either on single or composite tasks. Besides achieving reliable acoustic task specification for LALMs, this also reveals that LALMs exhibit certain "functional pathways" in their attention heads.
title AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2509.01787