CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908606573576192 |
|---|---|
| author | Xu, Guixian Su, Zeli Zhang, Ziyin Liu, Jianing Han, XU Zhang, Ting Dong, Yushuang |
| author_facet | Xu, Guixian Su, Zeli Zhang, Ziyin Liu, Jianing Han, XU Zhang, Ting Dong, Yushuang |
| contents | Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline generation. To address this gap, we introduce a novel dataset, Chinese Minority Headline Generation (CMHG), which includes 100,000 entries for Tibetan, and 50,000 entries each for Uyghur and Mongolian, specifically curated for headline generation tasks. Additionally, we propose a high-quality test set annotated by native speakers, designed to serve as a benchmark for future research in this domain. We hope this dataset will become a valuable resource for advancing headline generation in Chinese minority languages and contribute to the development of related benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_09990 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China Xu, Guixian Su, Zeli Zhang, Ziyin Liu, Jianing Han, XU Zhang, Ting Dong, Yushuang Computation and Language Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline generation. To address this gap, we introduce a novel dataset, Chinese Minority Headline Generation (CMHG), which includes 100,000 entries for Tibetan, and 50,000 entries each for Uyghur and Mongolian, specifically curated for headline generation tasks. Additionally, we propose a high-quality test set annotated by native speakers, designed to serve as a benchmark for future research in this domain. We hope this dataset will become a valuable resource for advancing headline generation in Chinese minority languages and contribute to the development of related benchmarks. |
| title | CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.09990 |