Keyword-Guided Adaptation of Automatic Speech Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shamsian, Aviv, Navon, Aviv, Glazer, Neta, Hetz, Gill, Keshet, Joseph
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917684832108544
author Shamsian, Aviv
Navon, Aviv
Glazer, Neta
Hetz, Gill
Keshet, Joseph
author_facet Shamsian, Aviv
Navon, Aviv
Glazer, Neta
Hetz, Gill
Keshet, Joseph
contents Automatic Speech Recognition (ASR) technology has made significant progress in recent years, providing accurate transcription across various domains. However, some challenges remain, especially in noisy environments and specialized jargon. In this paper, we propose a novel approach for improved jargon word recognition by contextual biasing Whisper-based models. We employ a keyword spotting model that leverages the Whisper encoder representation to dynamically generate prompts for guiding the decoder during the transcription process. We introduce two approaches to effectively steer the decoder towards these prompts: KG-Whisper, which is aimed at fine-tuning the Whisper decoder, and KG-Whisper-PT, which learns a prompt prefix. Our results show a significant improvement in the recognition accuracy of specified keywords and in reducing the overall word error rates. Specifically, in unseen language generalization, we demonstrate an average WER improvement of 5.1% over Whisper.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02649
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Keyword-Guided Adaptation of Automatic Speech Recognition
Shamsian, Aviv
Navon, Aviv
Glazer, Neta
Hetz, Gill
Keshet, Joseph
Audio and Speech Processing
Machine Learning
Sound
Automatic Speech Recognition (ASR) technology has made significant progress in recent years, providing accurate transcription across various domains. However, some challenges remain, especially in noisy environments and specialized jargon. In this paper, we propose a novel approach for improved jargon word recognition by contextual biasing Whisper-based models. We employ a keyword spotting model that leverages the Whisper encoder representation to dynamically generate prompts for guiding the decoder during the transcription process. We introduce two approaches to effectively steer the decoder towards these prompts: KG-Whisper, which is aimed at fine-tuning the Whisper decoder, and KG-Whisper-PT, which learns a prompt prefix. Our results show a significant improvement in the recognition accuracy of specified keywords and in reducing the overall word error rates. Specifically, in unseen language generalization, we demonstrate an average WER improvement of 5.1% over Whisper.
title Keyword-Guided Adaptation of Automatic Speech Recognition
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2406.02649