Text-only adaptation in LLM-based ASR through text denoising
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912962795536384 |
|---|---|
| author | Carofilis, Andrés Burdisso, Sergio Villatoro-Tello, Esaú Kumar, Shashi Hacioglu, Kadri Madikeri, Srikanth Rangappa, Pradeep E, Manjunath K Motlicek, Petr Venkatesan, Shankar Stolcke, Andreas |
| author_facet | Carofilis, Andrés Burdisso, Sergio Villatoro-Tello, Esaú Kumar, Shashi Hacioglu, Kadri Madikeri, Srikanth Rangappa, Pradeep E, Manjunath K Motlicek, Petr Venkatesan, Shankar Stolcke, Andreas |
| contents | Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on the target domain text often disrupts the critical alignment between the speech and text modality learned by the projector, degrading performance. We introduce a novel text-only adaptation method that frames this process as a text denoising task. Our approach trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_20900 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Text-only adaptation in LLM-based ASR through text denoising Carofilis, Andrés Burdisso, Sergio Villatoro-Tello, Esaú Kumar, Shashi Hacioglu, Kadri Madikeri, Srikanth Rangappa, Pradeep E, Manjunath K Motlicek, Petr Venkatesan, Shankar Stolcke, Andreas Sound Computation and Language Machine Learning Audio and Speech Processing Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on the target domain text often disrupts the critical alignment between the speech and text modality learned by the projector, degrading performance. We introduce a novel text-only adaptation method that frames this process as a text denoising task. Our approach trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods. |
| title | Text-only adaptation in LLM-based ASR through text denoising |
| topic | Sound Computation and Language Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2601.20900 |