Closing the Modality Reasoning Gap for Speech Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chaoren, Lu, Heng, Zhang, Xueyao, Liu, Shujie, Lu, Yan, Li, Jinyu, Wu, Zhizheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910148176379904
author Wang, Chaoren
Lu, Heng
Zhang, Xueyao
Liu, Shujie
Lu, Yan
Li, Jinyu
Wu, Zhizheng
author_facet Wang, Chaoren
Lu, Heng
Zhang, Xueyao
Liu, Shujie
Lu, Yan
Li, Jinyu
Wu, Zhizheng
contents Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech- and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions. Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Closing the Modality Reasoning Gap for Speech Large Language Models
Wang, Chaoren
Lu, Heng
Zhang, Xueyao
Liu, Shujie
Lu, Yan
Li, Jinyu
Wu, Zhizheng
Computation and Language
Sound
Audio and Speech Processing
Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech- and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions. Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
title Closing the Modality Reasoning Gap for Speech Large Language Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.05543