Saved in:
Bibliographic Details
Main Authors: Trinh, Viet Anh, He, Xinlu, Whitehill, Jacob
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.10779
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Classroom speech and lectures often contain named entities (NEs) such as names of people and special terminology. While automatic speech recognition (ASR) systems have achieved remarkable performance on general speech, the word error rate (WER) of state-of-the-art ASR remains high for named entities. Since NE are often the most critical keywords, misrecognizing them can affect all downstream applications, especially when the ASR functions as the front end of a complex system. In this paper, we introduce a large language model (LLM) revision pipeline to revise incorrect NEs in ASR predictions by leveraging not only the LLM's world knowledge and reasoning ability but also the available phonetic and semantic context. We also introduce the NER-MIT-OpenCourseWare dataset, containing 45 hours of data from MIT courses for development and testing. On this dataset, our proposed technique achieves up to 30\% relative WER reduction for NEs.