PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liang, Junhong, Zhang, Bojun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908550533480448
author Liang, Junhong
Zhang, Bojun
author_facet Liang, Junhong
Zhang, Bojun
contents Existing Chinese ASR correction methods have not effectively utilized Pinyin information, a unique feature of the Chinese language. In this study, we address this gap by proposing a \textbf{P}inyin \textbf{E}nhanced \textbf{R}ephrasing \textbf{L}anguage model (PERL) pipeline, designed explicitly for N-best correction scenarios. We conduct experiments on the Aishell-1 dataset and our newly proposed DoAD dataset. The results show that our approach outperforms baseline methods, achieving a 29.11\% reduction in Character Error Rate on Aishell-1 and around 70\% CER reduction on domain-specific datasets. PERL predicts the correct length of the output, leveraging the Pinyin information, which is embedded with a semantic model to perform phonetically similar corrections. Extensive experiments demonstrate the effectiveness of correcting wrong characters using N-best output and the low latency of our model.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03230
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction
Liang, Junhong
Zhang, Bojun
Computation and Language
Existing Chinese ASR correction methods have not effectively utilized Pinyin information, a unique feature of the Chinese language. In this study, we address this gap by proposing a \textbf{P}inyin \textbf{E}nhanced \textbf{R}ephrasing \textbf{L}anguage model (PERL) pipeline, designed explicitly for N-best correction scenarios. We conduct experiments on the Aishell-1 dataset and our newly proposed DoAD dataset. The results show that our approach outperforms baseline methods, achieving a 29.11\% reduction in Character Error Rate on Aishell-1 and around 70\% CER reduction on domain-specific datasets. PERL predicts the correct length of the output, leveraging the Pinyin information, which is embedded with a semantic model to perform phonetically similar corrections. Extensive experiments demonstrate the effectiveness of correcting wrong characters using N-best output and the low latency of our model.
title PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction
topic Computation and Language
url https://arxiv.org/abs/2412.03230