Unifying Global and Near-Context Biasing in a Single Trie Pass

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thorbecke, Iuliia, Villatoro-Tello, Esaú, Zuluaga-Gomez, Juan, Kumar, Shashi, Burdisso, Sergio, Rangappa, Pradeep, Carofilis, Andrés, Madikeri, Srikanth, Motlicek, Petr, Pandia, Karthik, Hacioğlu, Kadri, Stolcke, Andreas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909769120350208
author Thorbecke, Iuliia
Villatoro-Tello, Esaú
Zuluaga-Gomez, Juan
Kumar, Shashi
Burdisso, Sergio
Rangappa, Pradeep
Carofilis, Andrés
Madikeri, Srikanth
Motlicek, Petr
Pandia, Karthik
Hacioğlu, Kadri
Stolcke, Andreas
author_facet Thorbecke, Iuliia
Villatoro-Tello, Esaú
Zuluaga-Gomez, Juan
Kumar, Shashi
Burdisso, Sergio
Rangappa, Pradeep
Carofilis, Andrés
Madikeri, Srikanth
Motlicek, Petr
Pandia, Karthik
Hacioğlu, Kadri
Stolcke, Andreas
contents Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and in adapting to new domains using only text data. This work presents a practical approach to address these challenges through an unexplored combination of an NE bias list and a word-level n-gram language model (LM). This solution balances simplicity and effectiveness, improving entities' recognition while maintaining or even enhancing overall ASR performance. We efficiently integrate this enriched biasing method into a transducer-based ASR system, enabling context adaptation with almost no computational overhead. We present our results on three datasets spanning four languages and compare them to state-of-the-art biasing strategies. We demonstrate that the proposed combination of keyword biasing and n-gram LM improves entity recognition by up to 32% relative and reduces overall WER by up to a 12% relative.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13514
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unifying Global and Near-Context Biasing in a Single Trie Pass
Thorbecke, Iuliia
Villatoro-Tello, Esaú
Zuluaga-Gomez, Juan
Kumar, Shashi
Burdisso, Sergio
Rangappa, Pradeep
Carofilis, Andrés
Madikeri, Srikanth
Motlicek, Petr
Pandia, Karthik
Hacioğlu, Kadri
Stolcke, Andreas
Computation and Language
Sound
Audio and Speech Processing
Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and in adapting to new domains using only text data. This work presents a practical approach to address these challenges through an unexplored combination of an NE bias list and a word-level n-gram language model (LM). This solution balances simplicity and effectiveness, improving entities' recognition while maintaining or even enhancing overall ASR performance. We efficiently integrate this enriched biasing method into a transducer-based ASR system, enabling context adaptation with almost no computational overhead. We present our results on three datasets spanning four languages and compare them to state-of-the-art biasing strategies. We demonstrate that the proposed combination of keyword biasing and n-gram LM improves entity recognition by up to 32% relative and reduces overall WER by up to a 12% relative.
title Unifying Global and Near-Context Biasing in a Single Trie Pass
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.13514