Text-only adaptation in LLM-based ASR through text denoising

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Carofilis, Andrés, Burdisso, Sergio, Villatoro-Tello, Esaú, Kumar, Shashi, Hacioglu, Kadri, Madikeri, Srikanth, Rangappa, Pradeep, E, Manjunath K, Motlicek, Petr, Venkatesan, Shankar, Stolcke, Andreas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912962795536384
author Carofilis, Andrés
Burdisso, Sergio
Villatoro-Tello, Esaú
Kumar, Shashi
Hacioglu, Kadri
Madikeri, Srikanth
Rangappa, Pradeep
E, Manjunath K
Motlicek, Petr
Venkatesan, Shankar
Stolcke, Andreas
author_facet Carofilis, Andrés
Burdisso, Sergio
Villatoro-Tello, Esaú
Kumar, Shashi
Hacioglu, Kadri
Madikeri, Srikanth
Rangappa, Pradeep
E, Manjunath K
Motlicek, Petr
Venkatesan, Shankar
Stolcke, Andreas
contents Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on the target domain text often disrupts the critical alignment between the speech and text modality learned by the projector, degrading performance. We introduce a novel text-only adaptation method that frames this process as a text denoising task. Our approach trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20900
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Text-only adaptation in LLM-based ASR through text denoising
Carofilis, Andrés
Burdisso, Sergio
Villatoro-Tello, Esaú
Kumar, Shashi
Hacioglu, Kadri
Madikeri, Srikanth
Rangappa, Pradeep
E, Manjunath K
Motlicek, Petr
Venkatesan, Shankar
Stolcke, Andreas
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on the target domain text often disrupts the critical alignment between the speech and text modality learned by the projector, degrading performance. We introduce a novel text-only adaptation method that frames this process as a text denoising task. Our approach trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods.
title Text-only adaptation in LLM-based ASR through text denoising
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2601.20900