Improving Grammatical Error Correction via Contextual Data Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yixuan, Wang, Baoxin, Liu, Yijun, Zhu, Qingfu, Wu, Dayong, Che, Wanxiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910500645765120
author Wang, Yixuan
Wang, Baoxin
Liu, Yijun
Zhu, Qingfu
Wu, Dayong
Che, Wanxiang
author_facet Wang, Yixuan
Wang, Baoxin
Liu, Yijun
Zhu, Qingfu
Wu, Dayong
Che, Wanxiang
contents Nowadays, data augmentation through synthetic data has been widely used in the field of Grammatical Error Correction (GEC) to alleviate the problem of data scarcity. However, these synthetic data are mainly used in the pre-training phase rather than the data-limited fine-tuning phase due to inconsistent error distribution and noisy labels. In this paper, we propose a synthetic data construction method based on contextual augmentation, which can ensure an efficient augmentation of the original data with a more consistent error distribution. Specifically, we combine rule-based substitution with model-based generation, using the generative model to generate a richer context for the extracted error patterns. Besides, we also propose a relabeling-based data cleaning method to mitigate the effects of noisy labels in synthetic data. Experiments on CoNLL14 and BEA19-Test show that our proposed augmentation method consistently and substantially outperforms strong baselines and achieves the state-of-the-art level with only a few synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17456
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Grammatical Error Correction via Contextual Data Augmentation
Wang, Yixuan
Wang, Baoxin
Liu, Yijun
Zhu, Qingfu
Wu, Dayong
Che, Wanxiang
Computation and Language
Artificial Intelligence
Nowadays, data augmentation through synthetic data has been widely used in the field of Grammatical Error Correction (GEC) to alleviate the problem of data scarcity. However, these synthetic data are mainly used in the pre-training phase rather than the data-limited fine-tuning phase due to inconsistent error distribution and noisy labels. In this paper, we propose a synthetic data construction method based on contextual augmentation, which can ensure an efficient augmentation of the original data with a more consistent error distribution. Specifically, we combine rule-based substitution with model-based generation, using the generative model to generate a richer context for the extracted error patterns. Besides, we also propose a relabeling-based data cleaning method to mitigate the effects of noisy labels in synthetic data. Experiments on CoNLL14 and BEA19-Test show that our proposed augmentation method consistently and substantially outperforms strong baselines and achieves the state-of-the-art level with only a few synthetic data.
title Improving Grammatical Error Correction via Contextual Data Augmentation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.17456