CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yong, Meng, Fandong, Zhou, Jie
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914805242134528
author Hu, Yong
Meng, Fandong
Zhou, Jie
author_facet Hu, Yong
Meng, Fandong
Zhou, Jie
contents In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform. Compared with existing CSC datasets aimed at Chinese learners, CSCD-NS is ten times larger in scale and exhibits a distinct error distribution, with a significantly higher proportion of word-level errors. To further enhance the data resource, we propose a novel method that simulates the input process through an input method, generating large-scale and high-quality pseudo data that closely resembles the actual error distribution and outperforms existing methods. Moreover, we investigate the performance of various models in this scenario, including large language models (LLMs), such as ChatGPT. The result indicates that generative models underperform BERT-like classification models due to strict length and pronunciation constraints. The high prevalence of word-level errors also makes CSC for native speakers challenging enough, leaving substantial room for improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2211_08788
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
Hu, Yong
Meng, Fandong
Zhou, Jie
Computation and Language
Artificial Intelligence
In this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40,000 samples from a Chinese social platform. Compared with existing CSC datasets aimed at Chinese learners, CSCD-NS is ten times larger in scale and exhibits a distinct error distribution, with a significantly higher proportion of word-level errors. To further enhance the data resource, we propose a novel method that simulates the input process through an input method, generating large-scale and high-quality pseudo data that closely resembles the actual error distribution and outperforms existing methods. Moreover, we investigate the performance of various models in this scenario, including large language models (LLMs), such as ChatGPT. The result indicates that generative models underperform BERT-like classification models due to strict length and pronunciation constraints. The high prevalence of word-level errors also makes CSC for native speakers challenging enough, leaving substantial room for improvement.
title CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2211.08788