SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wei-Ping, Lin, Guan-Ting, Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915340588417024
author Huang, Wei-Ping
Lin, Guan-Ting
Lee, Hung-yi
author_facet Huang, Wei-Ping
Lin, Guan-Ting
Lee, Hung-yi
contents Despite progress in end-to-end ASR, real-world domain mismatches still cause performance drops, which Test-Time Adaptation (TTA) aims to mitigate by adjusting models during inference. Recent work explores combining TTA with external language models, using techniques like beam search rescoring or generative error correction. In this work, we identify a previously overlooked challenge: TTA can interfere with language model rescoring, revealing the nontrivial nature of effectively combining the two methods. Based on this insight, we propose SUTA-LM, a simple yet effective extension of SUTA, an entropy-minimization-based TTA approach, with language model rescoring. SUTA-LM first applies a controlled adaptation process guided by an auto-step selection mechanism leveraging both acoustic and linguistic information, followed by language model rescoring to refine the outputs. Experiments on 18 diverse ASR datasets show that SUTA-LM achieves robust results across a wide range of domains.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11121
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
Huang, Wei-Ping
Lin, Guan-Ting
Lee, Hung-yi
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Despite progress in end-to-end ASR, real-world domain mismatches still cause performance drops, which Test-Time Adaptation (TTA) aims to mitigate by adjusting models during inference. Recent work explores combining TTA with external language models, using techniques like beam search rescoring or generative error correction. In this work, we identify a previously overlooked challenge: TTA can interfere with language model rescoring, revealing the nontrivial nature of effectively combining the two methods. Based on this insight, we propose SUTA-LM, a simple yet effective extension of SUTA, an entropy-minimization-based TTA approach, with language model rescoring. SUTA-LM first applies a controlled adaptation process guided by an auto-step selection mechanism leveraging both acoustic and linguistic information, followed by language model rescoring to refine the outputs. Experiments on 18 diverse ASR datasets show that SUTA-LM achieves robust results across a wide range of domains.
title SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.11121