Error-Robust Retrieval for Chinese Spelling Check

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Xunjian, Hu, Xinyu, Jiang, Jin, Wan, Xiaojun
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909118581702656
author Yin, Xunjian
Hu, Xinyu
Jiang, Jin
Wan, Xiaojun
author_facet Yin, Xunjian
Hu, Xinyu
Jiang, Jin
Wan, Xiaojun
contents Chinese Spelling Check (CSC) aims to detect and correct error tokens in Chinese contexts, which has a wide range of applications. However, it is confronted with the challenges of insufficient annotated data and the issue that previous methods may actually not fully leverage the existing datasets. In this paper, we introduce our plug-and-play retrieval method with error-robust information for Chinese Spelling Check (RERIC), which can be directly applied to existing CSC models. The datastore for retrieval is built completely based on the training data, with elaborate designs according to the characteristics of CSC. Specifically, we employ multimodal representations that fuse phonetic, morphologic, and contextual information in the calculation of query and key during retrieval to enhance robustness against potential errors. Furthermore, in order to better judge the retrieved candidates, the n-gram surrounding the token to be checked is regarded as the value and utilized for specific reranking. The experiment results on the SIGHAN benchmarks demonstrate that our proposed method achieves substantial improvements over existing work.
format Preprint
id arxiv_https___arxiv_org_abs_2211_07843
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Error-Robust Retrieval for Chinese Spelling Check
Yin, Xunjian
Hu, Xinyu
Jiang, Jin
Wan, Xiaojun
Computation and Language
Chinese Spelling Check (CSC) aims to detect and correct error tokens in Chinese contexts, which has a wide range of applications. However, it is confronted with the challenges of insufficient annotated data and the issue that previous methods may actually not fully leverage the existing datasets. In this paper, we introduce our plug-and-play retrieval method with error-robust information for Chinese Spelling Check (RERIC), which can be directly applied to existing CSC models. The datastore for retrieval is built completely based on the training data, with elaborate designs according to the characteristics of CSC. Specifically, we employ multimodal representations that fuse phonetic, morphologic, and contextual information in the calculation of query and key during retrieval to enhance robustness against potential errors. Furthermore, in order to better judge the retrieved candidates, the n-gram surrounding the token to be checked is regarded as the value and utilized for specific reranking. The experiment results on the SIGHAN benchmarks demonstrate that our proposed method achieves substantial improvements over existing work.
title Error-Robust Retrieval for Chinese Spelling Check
topic Computation and Language
url https://arxiv.org/abs/2211.07843