CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Guixian, Su, Zeli, Zhang, Ziyin, Liu, Jianing, Han, XU, Zhang, Ting, Dong, Yushuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908606573576192
author Xu, Guixian
Su, Zeli
Zhang, Ziyin
Liu, Jianing
Han, XU
Zhang, Ting
Dong, Yushuang
author_facet Xu, Guixian
Su, Zeli
Zhang, Ziyin
Liu, Jianing
Han, XU
Zhang, Ting
Dong, Yushuang
contents Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline generation. To address this gap, we introduce a novel dataset, Chinese Minority Headline Generation (CMHG), which includes 100,000 entries for Tibetan, and 50,000 entries each for Uyghur and Mongolian, specifically curated for headline generation tasks. Additionally, we propose a high-quality test set annotated by native speakers, designed to serve as a benchmark for future research in this domain. We hope this dataset will become a valuable resource for advancing headline generation in Chinese minority languages and contribute to the development of related benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
Xu, Guixian
Su, Zeli
Zhang, Ziyin
Liu, Jianing
Han, XU
Zhang, Ting
Dong, Yushuang
Computation and Language
Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline generation. To address this gap, we introduce a novel dataset, Chinese Minority Headline Generation (CMHG), which includes 100,000 entries for Tibetan, and 50,000 entries each for Uyghur and Mongolian, specifically curated for headline generation tasks. Additionally, we propose a high-quality test set annotated by native speakers, designed to serve as a benchmark for future research in this domain. We hope this dataset will become a valuable resource for advancing headline generation in Chinese minority languages and contribute to the development of related benchmarks.
title CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
topic Computation and Language
url https://arxiv.org/abs/2509.09990