Unifying Global and Near-Context Biasing in a Single Trie Pass
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909769120350208 |
|---|---|
| author | Thorbecke, Iuliia Villatoro-Tello, Esaú Zuluaga-Gomez, Juan Kumar, Shashi Burdisso, Sergio Rangappa, Pradeep Carofilis, Andrés Madikeri, Srikanth Motlicek, Petr Pandia, Karthik Hacioğlu, Kadri Stolcke, Andreas |
| author_facet | Thorbecke, Iuliia Villatoro-Tello, Esaú Zuluaga-Gomez, Juan Kumar, Shashi Burdisso, Sergio Rangappa, Pradeep Carofilis, Andrés Madikeri, Srikanth Motlicek, Petr Pandia, Karthik Hacioğlu, Kadri Stolcke, Andreas |
| contents | Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and in adapting to new domains using only text data. This work presents a practical approach to address these challenges through an unexplored combination of an NE bias list and a word-level n-gram language model (LM). This solution balances simplicity and effectiveness, improving entities' recognition while maintaining or even enhancing overall ASR performance. We efficiently integrate this enriched biasing method into a transducer-based ASR system, enabling context adaptation with almost no computational overhead. We present our results on three datasets spanning four languages and compare them to state-of-the-art biasing strategies. We demonstrate that the proposed combination of keyword biasing and n-gram LM improves entity recognition by up to 32% relative and reduces overall WER by up to a 12% relative. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_13514 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Unifying Global and Near-Context Biasing in a Single Trie Pass Thorbecke, Iuliia Villatoro-Tello, Esaú Zuluaga-Gomez, Juan Kumar, Shashi Burdisso, Sergio Rangappa, Pradeep Carofilis, Andrés Madikeri, Srikanth Motlicek, Petr Pandia, Karthik Hacioğlu, Kadri Stolcke, Andreas Computation and Language Sound Audio and Speech Processing Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and in adapting to new domains using only text data. This work presents a practical approach to address these challenges through an unexplored combination of an NE bias list and a word-level n-gram language model (LM). This solution balances simplicity and effectiveness, improving entities' recognition while maintaining or even enhancing overall ASR performance. We efficiently integrate this enriched biasing method into a transducer-based ASR system, enabling context adaptation with almost no computational overhead. We present our results on three datasets spanning four languages and compare them to state-of-the-art biasing strategies. We demonstrate that the proposed combination of keyword biasing and n-gram LM improves entity recognition by up to 32% relative and reduces overall WER by up to a 12% relative. |
| title | Unifying Global and Near-Context Biasing in a Single Trie Pass |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.13514 |