Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Zhennan, Huang, Kaixun, Ren, Wei, Yang, Linju, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910973641621504
author Lin, Zhennan
Huang, Kaixun
Ren, Wei
Yang, Linju
Xie, Lei
author_facet Lin, Zhennan
Huang, Kaixun
Ren, Wei
Yang, Linju
Xie, Lei
contents Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromising contextual phrase integrity, leading to accuracy reduction. In this paper, we propose an encoder-based phrase-level contextualized ASR method that leverages dynamic vocabulary prediction and activation. We introduce architectural optimizations and integrate a bias loss to extend phrase-level predictions based on frame-level outputs. We also introduce a confidence-activated decoding method that ensures the complete output of contextual phrases while suppressing incorrect bias. Experiments on Librispeech and Wenetspeech datasets demonstrate that our approach achieves relative WER reductions of 28.31% and 23.49% compared to baseline, with the WER on contextual phrases decreasing relatively by 72.04% and 75.69%.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
Lin, Zhennan
Huang, Kaixun
Ren, Wei
Yang, Linju
Xie, Lei
Sound
Computation and Language
Audio and Speech Processing
Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromising contextual phrase integrity, leading to accuracy reduction. In this paper, we propose an encoder-based phrase-level contextualized ASR method that leverages dynamic vocabulary prediction and activation. We introduce architectural optimizations and integrate a bias loss to extend phrase-level predictions based on frame-level outputs. We also introduce a confidence-activated decoding method that ensures the complete output of contextual phrases while suppressing incorrect bias. Experiments on Librispeech and Wenetspeech datasets demonstrate that our approach achieves relative WER reductions of 28.31% and 23.49% compared to baseline, with the WER on contextual phrases decreasing relatively by 72.04% and 75.69%.
title Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2505.23077