Peeking Into The Future For Contextual Biasing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Selvakumar, Ramaneswaran, Tseng, Cindy, Kim, Eesung, Apsingekar, Vijendra Raj, Tang, Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912777040297984
author Selvakumar, Ramaneswaran
Tseng, Cindy
Kim, Eesung
Apsingekar, Vijendra Raj
Tang, Yun
author_facet Selvakumar, Ramaneswaran
Tseng, Cindy
Kim, Eesung
Apsingekar, Vijendra Raj
Tang, Yun
contents While end-to-end (E2E) automatic speech recognition (ASR) models excel at general transcription, they struggle to recognize rare or unseen named entities (e.g., contact names, locations), which are critical for downstream applications like virtual assistants. In this paper, we propose a contextual biasing method for attention based encoder decoder (AED) models using a list of candidate named entities. Instead of predicting only the next token, we simultaneously predict multiple future tokens, enabling the model to "peek into the future" and score potential candidate entities in the entity list. Moreover, our approach leverages the multi-token prediction logits directly without requiring additional entity encoders or cross-attention layers, significantly reducing architectural complexity. Experiments on Librispeech demonstrate that our approach achieves up to 50.34% relative improvement in named entity word error rate compared to the baseline AED model.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17657
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Peeking Into The Future For Contextual Biasing
Selvakumar, Ramaneswaran
Tseng, Cindy
Kim, Eesung
Apsingekar, Vijendra Raj
Tang, Yun
Computation and Language
While end-to-end (E2E) automatic speech recognition (ASR) models excel at general transcription, they struggle to recognize rare or unseen named entities (e.g., contact names, locations), which are critical for downstream applications like virtual assistants. In this paper, we propose a contextual biasing method for attention based encoder decoder (AED) models using a list of candidate named entities. Instead of predicting only the next token, we simultaneously predict multiple future tokens, enabling the model to "peek into the future" and score potential candidate entities in the entity list. Moreover, our approach leverages the multi-token prediction logits directly without requiring additional entity encoders or cross-attention layers, significantly reducing architectural complexity. Experiments on Librispeech demonstrate that our approach achieves up to 50.34% relative improvement in named entity word error rate compared to the baseline AED model.
title Peeking Into The Future For Contextual Biasing
topic Computation and Language
url https://arxiv.org/abs/2512.17657