Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Changsong, Peng, Yizhou, Chng, Eng Siong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912565677785088
author Liu, Changsong
Peng, Yizhou
Chng, Eng Siong
author_facet Liu, Changsong
Peng, Yizhou
Chng, Eng Siong
contents Contextual automatic speech recognition (ASR) systems allow for recognizing out-of-vocabulary (OOV) words, such as named entities or rare words. However, it remains challenging due to limited training data and ambiguous or inconsistent pronunciations. In this paper, we propose a synthesis-driven multi-pronunciation contextual biasing method that performs zero-shot contextual ASR on a pretrained Whisper model. Specifically, we leverage text-to-speech (TTS) systems to synthesize diverse speech samples containing each target rare word, and then use the pretrained Whisper model to extract multiple predicted pronunciation variants. These variant token sequences are compiled into a prefix-trie, which assigns rewards to beam hypotheses in a shallow-fusion manner during beam-search decoding. Subsequently, any recognized variant is mapped back to the original rare word in the final transcription. The evaluation results on the LibriSpeech dataset show that our method reduces biased-word error rate (B-WER) by 43% on test-clean and 44% on test-other while maintaining unbiased-WER (U-WER) essentially unchanged.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17796
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
Liu, Changsong
Peng, Yizhou
Chng, Eng Siong
Computation and Language
Audio and Speech Processing
Contextual automatic speech recognition (ASR) systems allow for recognizing out-of-vocabulary (OOV) words, such as named entities or rare words. However, it remains challenging due to limited training data and ambiguous or inconsistent pronunciations. In this paper, we propose a synthesis-driven multi-pronunciation contextual biasing method that performs zero-shot contextual ASR on a pretrained Whisper model. Specifically, we leverage text-to-speech (TTS) systems to synthesize diverse speech samples containing each target rare word, and then use the pretrained Whisper model to extract multiple predicted pronunciation variants. These variant token sequences are compiled into a prefix-trie, which assigns rewards to beam hypotheses in a shallow-fusion manner during beam-search decoding. Subsequently, any recognized variant is mapped back to the original rare word in the final transcription. The evaluation results on the LibriSpeech dataset show that our method reduces biased-word error rate (B-WER) by 43% on test-clean and 44% on test-other while maintaining unbiased-WER (U-WER) essentially unchanged.
title Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2508.17796