Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Jielin, Zhu, Ming, Zhao, Wenting, Liu, Zhiwei, Yang, Liangwei, Chen, Zixiang, Ram, Roshan, Prabhakar, Akshara, Tan, Juntao, Murthy, Rithesh, Heinecke, Shelby, Xiong, Caiming, Savarese, Silvio, Wang, Huan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914466400043008
author Qiu, Jielin
Zhu, Ming
Zhao, Wenting
Liu, Zhiwei
Yang, Liangwei
Chen, Zixiang
Ram, Roshan
Prabhakar, Akshara
Tan, Juntao
Murthy, Rithesh
Heinecke, Shelby
Xiong, Caiming
Savarese, Silvio
Wang, Huan
author_facet Qiu, Jielin
Zhu, Ming
Zhao, Wenting
Liu, Zhiwei
Yang, Liangwei
Chen, Zixiang
Ram, Roshan
Prabhakar, Akshara
Tan, Juntao
Murthy, Rithesh
Heinecke, Shelby
Xiong, Caiming
Savarese, Silvio
Wang, Huan
contents Audio-native large language models (audio-LLMs) commonly use Whisper as their audio encoder. However, Whisper was trained exclusively on speech data, producing weak representations for music and environmental sound. This forces downstream audio-LLMs to compensate through extensive training on large-scale non-speech data. We present Whisper-AuT, a domain-adapted audio encoder obtained by fine-tuning Whisper-large-v3 on a curated mixture of speech (80%), environmental sound (10%), and music (10%) totaling approximately 20M samples. The full encoder-decoder is trained end-to-end with a seq2seq captioning objective; the decoder is then discarded and only the encoder is retained. Linear probe evaluations show that Whisper-AuT achieves +23.0% on ESC-50 (environmental sound), +5.0% on GTZAN (music genre), and +0.7% on Speech Commands (keyword spotting) compared to the original Whisperlarge-v3 encoder. Whisper-AuT is designed as a drop-in replacement for Whisper in audio-LLM architectures, with the goal of reducing downstream training cost by providing stronger initial audio representations for non-speech domains.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10438
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
Qiu, Jielin
Zhu, Ming
Zhao, Wenting
Liu, Zhiwei
Yang, Liangwei
Chen, Zixiang
Ram, Roshan
Prabhakar, Akshara
Tan, Juntao
Murthy, Rithesh
Heinecke, Shelby
Xiong, Caiming
Savarese, Silvio
Wang, Huan
Sound
Audio-native large language models (audio-LLMs) commonly use Whisper as their audio encoder. However, Whisper was trained exclusively on speech data, producing weak representations for music and environmental sound. This forces downstream audio-LLMs to compensate through extensive training on large-scale non-speech data. We present Whisper-AuT, a domain-adapted audio encoder obtained by fine-tuning Whisper-large-v3 on a curated mixture of speech (80%), environmental sound (10%), and music (10%) totaling approximately 20M samples. The full encoder-decoder is trained end-to-end with a seq2seq captioning objective; the decoder is then discarded and only the encoder is retained. Linear probe evaluations show that Whisper-AuT achieves +23.0% on ESC-50 (environmental sound), +5.0% on GTZAN (music genre), and +0.7% on Speech Commands (keyword spotting) compared to the original Whisperlarge-v3 encoder. Whisper-AuT is designed as a drop-in replacement for Whisper in audio-LLM architectures, with the goal of reducing downstream training cost by providing stronger initial audio representations for non-speech domains.
title Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
topic Sound
url https://arxiv.org/abs/2604.10438