Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chao, Zheng, Rui-Chen, Ai, Yang, Ling, Zhen-Hua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912612703272960
author Wang, Chao
Zheng, Rui-Chen
Ai, Yang
Ling, Zhen-Hua
author_facet Wang, Chao
Zheng, Rui-Chen
Ai, Yang
Ling, Zhen-Hua
contents The integration of speech into Large Language Models (LLMs) has substantially expanded their capabilities, but often at the cost of weakening their core textual competence. This degradation limits the ability of speech-enabled LLMs to fully exploit their pre-trained text-based knowledge. In this work, we analyze the underlying mechanisms of this issue through a focused study of the widely used encoder-adaptor paradigm. We propose an analytical framework based on parameter importance estimation, which reveals that fine-tuning for speech introduces a textual importance distribution shift: the layer-wise allocation of parameters critical to textual reasoning is disrupted. Building on this insight, we investigate two mitigation strategies: layer-wise learning rate scheduling and Low-Rank Adaptation (LoRA), both aim to preserve the original parameter distribution. Experimental results show that both approaches better maintain textual competence than full fine-tuning, while also improving downstream spoken question answering performance. Furthermore, our analysis offers a principled explanation for the effectiveness of the proposed mitigation strategies, linking their benefits to the structural properties of textual knowledge in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23755
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
Wang, Chao
Zheng, Rui-Chen
Ai, Yang
Ling, Zhen-Hua
Computation and Language
Artificial Intelligence
The integration of speech into Large Language Models (LLMs) has substantially expanded their capabilities, but often at the cost of weakening their core textual competence. This degradation limits the ability of speech-enabled LLMs to fully exploit their pre-trained text-based knowledge. In this work, we analyze the underlying mechanisms of this issue through a focused study of the widely used encoder-adaptor paradigm. We propose an analytical framework based on parameter importance estimation, which reveals that fine-tuning for speech introduces a textual importance distribution shift: the layer-wise allocation of parameters critical to textual reasoning is disrupted. Building on this insight, we investigate two mitigation strategies: layer-wise learning rate scheduling and Low-Rank Adaptation (LoRA), both aim to preserve the original parameter distribution. Experimental results show that both approaches better maintain textual competence than full fine-tuning, while also improving downstream spoken question answering performance. Furthermore, our analysis offers a principled explanation for the effectiveness of the proposed mitigation strategies, linking their benefits to the structural properties of textual knowledge in LLMs.
title Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.23755