Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Yangui, Peng, Jing, Li, Xu, Xi, Yu, Zhang, Chengwei, Zhong, Guohui, Yu, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915691975671808
author Fang, Yangui
Peng, Jing
Li, Xu
Xi, Yu
Zhang, Chengwei
Zhong, Guohui
Yu, Kai
author_facet Fang, Yangui
Peng, Jing
Li, Xu
Xi, Yu
Zhang, Chengwei
Zhong, Guohui
Yu, Kai
contents Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performance. However, adapting them to new domains remains challenging, especially in low-resource settings where paired speech-text data is scarce. We propose a text-only fine-tuning strategy for Speech LLMs using unpaired target-domain text without requiring additional audio. To preserve speech-text alignment, we introduce a real-time evaluation mechanism during fine-tuning. This enables effective domain adaptation while maintaining source-domain performance. Experiments on LibriSpeech, SlideSpeech, and Medical datasets show that our method achieves competitive recognition performance, with minimal degradation compared to full audio-text fine-tuning. It also improves generalization to new domains without catastrophic forgetting, highlighting the potential of text-only fine-tuning for low-resource domain adaptation of ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05671
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
Fang, Yangui
Peng, Jing
Li, Xu
Xi, Yu
Zhang, Chengwei
Zhong, Guohui
Yu, Kai
Audio and Speech Processing
Computation and Language
Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performance. However, adapting them to new domains remains challenging, especially in low-resource settings where paired speech-text data is scarce. We propose a text-only fine-tuning strategy for Speech LLMs using unpaired target-domain text without requiring additional audio. To preserve speech-text alignment, we introduce a real-time evaluation mechanism during fine-tuning. This enables effective domain adaptation while maintaining source-domain performance. Experiments on LibriSpeech, SlideSpeech, and Medical datasets show that our method achieves competitive recognition performance, with minimal degradation compared to full audio-text fine-tuning. It also improves generalization to new domains without catastrophic forgetting, highlighting the potential of text-only fine-tuning for low-resource domain adaptation of ASR.
title Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2506.05671