Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Yuchun, Hu, Bo, He, Yifeng, Shi, Hao, Wang, Longbiao, Dang, Jianwu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909260536872960
author Shu, Yuchun
Hu, Bo
He, Yifeng
Shi, Hao
Wang, Longbiao
Dang, Jianwu
author_facet Shu, Yuchun
Hu, Bo
He, Yifeng
Shi, Hao
Wang, Longbiao
Dang, Jianwu
contents Accurately finding the wrong words in the automatic speech recognition (ASR) hypothesis and recovering them well-founded is the goal of speech error correction. In this paper, we propose a non-autoregressive speech error correction method. A Confidence Module measures the uncertainty of each word of the N-best ASR hypotheses as the reference to find the wrong word position. Besides, the acoustic feature from the ASR encoder is also used to provide the correct pronunciation references. N-best candidates from ASR are aligned using the edit path, to confirm each other and recover some missing character errors. Furthermore, the cross-attention mechanism fuses the information between error correction references and the ASR hypothesis. The experimental results show that both the acoustic and confidence references help with error correction. The proposed system reduces the error rate by 21% compared with the ASR model.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12817
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
Shu, Yuchun
Hu, Bo
He, Yifeng
Shi, Hao
Wang, Longbiao
Dang, Jianwu
Computation and Language
Sound
Audio and Speech Processing
Accurately finding the wrong words in the automatic speech recognition (ASR) hypothesis and recovering them well-founded is the goal of speech error correction. In this paper, we propose a non-autoregressive speech error correction method. A Confidence Module measures the uncertainty of each word of the N-best ASR hypotheses as the reference to find the wrong word position. Besides, the acoustic feature from the ASR encoder is also used to provide the correct pronunciation references. N-best candidates from ASR are aligned using the edit path, to confirm each other and recover some missing character errors. Furthermore, the cross-attention mechanism fuses the information between error correction references and the ASR hypothesis. The experimental results show that both the acoustic and confidence references help with error correction. The proposed system reduces the error rate by 21% compared with the ASR model.
title Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.12817