Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Ruitao, Zhang, Bixi, Liang, Sheng, Yuan, Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914095068872704
author Feng, Ruitao
Zhang, Bixi
Liang, Sheng
Yuan, Zheng
author_facet Feng, Ruitao
Zhang, Bixi
Liang, Sheng
Yuan, Zheng
contents Aligning pretrained audio encoders and Large Language Models (LLMs) offers a promising, parameter-efficient path to building powerful multimodal agents. However, existing methods often require costly full-model finetuning or rely on static adapters that may lack expressive power. Drawing inspiration from the Platonic Representation Hypothesis, we introduce SteerMoE, a novel and modular framework for audio-language alignment. SteerMoE freezes both the audio encoder and the LLM decoder, training only a lightweight steering module integrated within the encoder's layers. This module uses a Mixture-of-Experts (MoE) router to dynamically select and apply learned steering vectors, progressively transforming continuous audio representations into a space comprehensible to the LLM. By operating entirely in the continuous embedding space, our approach requires no modifications to the LLM's vocabulary and preserves its advanced reasoning and agentic capabilities. We demonstrate through experiments on ASR, audio understanding, and a qualitative function-calling task that SteerMoE achieves strong performance while remaining highly modular and computationally efficient, offering a robust new paradigm for developing sophisticated audio-language systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
Feng, Ruitao
Zhang, Bixi
Liang, Sheng
Yuan, Zheng
Sound
I.2.7
Aligning pretrained audio encoders and Large Language Models (LLMs) offers a promising, parameter-efficient path to building powerful multimodal agents. However, existing methods often require costly full-model finetuning or rely on static adapters that may lack expressive power. Drawing inspiration from the Platonic Representation Hypothesis, we introduce SteerMoE, a novel and modular framework for audio-language alignment. SteerMoE freezes both the audio encoder and the LLM decoder, training only a lightweight steering module integrated within the encoder's layers. This module uses a Mixture-of-Experts (MoE) router to dynamically select and apply learned steering vectors, progressively transforming continuous audio representations into a space comprehensible to the LLM. By operating entirely in the continuous embedding space, our approach requires no modifications to the LLM's vocabulary and preserves its advanced reasoning and agentic capabilities. We demonstrate through experiments on ASR, audio understanding, and a qualitative function-calling task that SteerMoE achieves strong performance while remaining highly modular and computationally efficient, offering a robust new paradigm for developing sophisticated audio-language systems.
title Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
topic Sound
I.2.7
url https://arxiv.org/abs/2510.13558