EdaCSC: Two Easy Data Augmentation Methods for Chinese Spelling Correction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sheng, Lei, Xu, Shuai-Shuai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912018855886848
author Sheng, Lei
Xu, Shuai-Shuai
author_facet Sheng, Lei
Xu, Shuai-Shuai
contents Chinese Spelling Correction (CSC) aims to detect and correct spelling errors in Chinese sentences caused by phonetic or visual similarities. While current CSC models integrate pinyin or glyph features and have shown significant progress,they still face challenges when dealing with sentences containing multiple typos and are susceptible to overcorrection in real-world scenarios. In contrast to existing model-centric approaches, we propose two data augmentation methods to address these limitations. Firstly, we augment the dataset by either splitting long sentences into shorter ones or reducing typos in sentences with multiple typos. Subsequently, we employ different training processes to select the optimal model. Experimental evaluations on the SIGHAN benchmarks demonstrate the superiority of our approach over most existing models, achieving state-of-the-art performance on the SIGHAN15 test set.
format Preprint
id arxiv_https___arxiv_org_abs_2409_05105
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EdaCSC: Two Easy Data Augmentation Methods for Chinese Spelling Correction
Sheng, Lei
Xu, Shuai-Shuai
Computation and Language
Artificial Intelligence
Chinese Spelling Correction (CSC) aims to detect and correct spelling errors in Chinese sentences caused by phonetic or visual similarities. While current CSC models integrate pinyin or glyph features and have shown significant progress,they still face challenges when dealing with sentences containing multiple typos and are susceptible to overcorrection in real-world scenarios. In contrast to existing model-centric approaches, we propose two data augmentation methods to address these limitations. Firstly, we augment the dataset by either splitting long sentences into shorter ones or reducing typos in sentences with multiple typos. Subsequently, we employ different training processes to select the optimal model. Experimental evaluations on the SIGHAN benchmarks demonstrate the superiority of our approach over most existing models, achieving state-of-the-art performance on the SIGHAN15 test set.
title EdaCSC: Two Easy Data Augmentation Methods for Chinese Spelling Correction
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.05105