Retrieval Augmented Correction of Named Entity Speech Recognition Errors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pusateri, Ernest, Walia, Anmol, Kashi, Anirudh, Bandyopadhyay, Bortik, Hyder, Nadia, Mahinder, Sayantan, Anantha, Raviteja, Liu, Daben, Gondala, Sashank
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910596413259776
author Pusateri, Ernest
Walia, Anmol
Kashi, Anirudh
Bandyopadhyay, Bortik
Hyder, Nadia
Mahinder, Sayantan
Anantha, Raviteja
Liu, Daben
Gondala, Sashank
author_facet Pusateri, Ernest
Walia, Anmol
Kashi, Anirudh
Bandyopadhyay, Bortik
Hyder, Nadia
Mahinder, Sayantan
Anantha, Raviteja
Liu, Daben
Gondala, Sashank
contents In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database of relevant knowledge is available, retrieval augmented generation (RAG) has achieved impressive results when used with LLMs. In this work, we propose a RAG-like technique for correcting speech recognition entity name errors. Our approach uses a vector database to index a set of relevant entities. At runtime, database queries are generated from possibly errorful textual ASR hypotheses, and the entities retrieved using these queries are fed, along with the ASR hypotheses, to an LLM which has been adapted to correct ASR errors. Overall, our best system achieves 33%-39% relative word error rate reductions on synthetic test sets focused on voice assistant queries of rare music entities without regressing on the STOP test set, a publicly available voice assistant test set covering many domains.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06062
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrieval Augmented Correction of Named Entity Speech Recognition Errors
Pusateri, Ernest
Walia, Anmol
Kashi, Anirudh
Bandyopadhyay, Bortik
Hyder, Nadia
Mahinder, Sayantan
Anantha, Raviteja
Liu, Daben
Gondala, Sashank
Audio and Speech Processing
Sound
In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database of relevant knowledge is available, retrieval augmented generation (RAG) has achieved impressive results when used with LLMs. In this work, we propose a RAG-like technique for correcting speech recognition entity name errors. Our approach uses a vector database to index a set of relevant entities. At runtime, database queries are generated from possibly errorful textual ASR hypotheses, and the entities retrieved using these queries are fed, along with the ASR hypotheses, to an LLM which has been adapted to correct ASR errors. Overall, our best system achieves 33%-39% relative word error rate reductions on synthetic test sets focused on voice assistant queries of rare music entities without regressing on the STOP test set, a publicly available voice assistant test set covering many domains.
title Retrieval Augmented Correction of Named Entity Speech Recognition Errors
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.06062