Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hu, Yuchen, Chen, Chen, Qin, Chengwei, Zhu, Qiushi, Chng, Eng Siong, Li, Ruizhe
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917667865100288
author Hu, Yuchen
Chen, Chen
Qin, Chengwei
Zhu, Qiushi
Chng, Eng Siong
Li, Ruizhe
author_facet Hu, Yuchen
Chen, Chen
Qin, Chengwei
Zhu, Qiushi
Chng, Eng Siong
Li, Ruizhe
contents Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which aims to predict the ground-truth transcription from the decoded N-best hypotheses. Thanks to the strong language generation ability of LLMs and rich information in the N-best list, GER shows great effectiveness in enhancing ASR results. However, it still suffers from two limitations: 1) LLMs are unaware of the source speech during GER, which may lead to results that are grammatically correct but violate the source speech content, 2) N-best hypotheses usually only vary in a few tokens, making it redundant to send all of them for GER, which could confuse LLM about which tokens to focus on and thus lead to increased miscorrection. In this paper, we propose ClozeGER, a new paradigm for ASR generative error correction. First, we introduce a multimodal LLM (i.e., SpeechGPT) to receive source speech as extra input to improve the fidelity of correction output. Then, we reformat GER as a cloze test with logits calibration to remove the input information redundancy and simplify GER with clear instructions. Experiments show that ClozeGER achieves a new breakthrough over vanilla GER on 9 popular ASR datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
Hu, Yuchen
Chen, Chen
Qin, Chengwei
Zhu, Qiushi
Chng, Eng Siong
Li, Ruizhe
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which aims to predict the ground-truth transcription from the decoded N-best hypotheses. Thanks to the strong language generation ability of LLMs and rich information in the N-best list, GER shows great effectiveness in enhancing ASR results. However, it still suffers from two limitations: 1) LLMs are unaware of the source speech during GER, which may lead to results that are grammatically correct but violate the source speech content, 2) N-best hypotheses usually only vary in a few tokens, making it redundant to send all of them for GER, which could confuse LLM about which tokens to focus on and thus lead to increased miscorrection. In this paper, we propose ClozeGER, a new paradigm for ASR generative error correction. First, we introduce a multimodal LLM (i.e., SpeechGPT) to receive source speech as extra input to improve the fidelity of correction output. Then, we reformat GER as a cloze test with logits calibration to remove the input information redundancy and simplify GER with clear instructions. Experiments show that ClozeGER achieves a new breakthrough over vanilla GER on 9 popular ASR datasets.
title Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.10025