MCTS: A Multi-Reference Chinese Text Simplification Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chong, Ruining, Lu, Luming, Yang, Liner, Nie, Jinran, Liu, Zhenghao, Wang, Shuo, Zhou, Shuhan, Li, Yaoxin, Yang, Erhong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913378423799808
author Chong, Ruining
Lu, Luming
Yang, Liner
Nie, Jinran
Liu, Zhenghao
Wang, Shuo
Zhou, Shuhan
Li, Yaoxin
Yang, Erhong
author_facet Chong, Ruining
Lu, Luming
Yang, Liner
Nie, Jinran
Liu, Zhenghao
Wang, Shuo
Zhou, Shuhan
Li, Yaoxin
Yang, Erhong
contents Text simplification aims to make the text easier to understand by applying rewriting transformations. There has been very little research on Chinese text simplification for a long time. The lack of generic evaluation data is an essential reason for this phenomenon. In this paper, we introduce MCTS, a multi-reference Chinese text simplification dataset. We describe the annotation process of the dataset and provide a detailed analysis. Furthermore, we evaluate the performance of several unsupervised methods and advanced large language models. We additionally provide Chinese text simplification parallel data that can be used for training, acquired by utilizing machine translation and English text simplification. We hope to build a basic understanding of Chinese text simplification through the foundational work and provide references for future research. All of the code and data are released at https://github.com/blcuicall/mcts/.
format Preprint
id arxiv_https___arxiv_org_abs_2306_02796
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MCTS: A Multi-Reference Chinese Text Simplification Dataset
Chong, Ruining
Lu, Luming
Yang, Liner
Nie, Jinran
Liu, Zhenghao
Wang, Shuo
Zhou, Shuhan
Li, Yaoxin
Yang, Erhong
Computation and Language
Text simplification aims to make the text easier to understand by applying rewriting transformations. There has been very little research on Chinese text simplification for a long time. The lack of generic evaluation data is an essential reason for this phenomenon. In this paper, we introduce MCTS, a multi-reference Chinese text simplification dataset. We describe the annotation process of the dataset and provide a detailed analysis. Furthermore, we evaluate the performance of several unsupervised methods and advanced large language models. We additionally provide Chinese text simplification parallel data that can be used for training, acquired by utilizing machine translation and English text simplification. We hope to build a basic understanding of Chinese text simplification through the foundational work and provide references for future research. All of the code and data are released at https://github.com/blcuicall/mcts/.
title MCTS: A Multi-Reference Chinese Text Simplification Dataset
topic Computation and Language
url https://arxiv.org/abs/2306.02796