CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chengwei, Wang, Jiapu, Gao, Mingyang, Zhuo, Xingrui, Guo, Jipeng, Lei, Runlin, Luo, Haoran, Chen, Tianyu, Zhou, Haoyi, Pan, Shirui, Li, Zechao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912633371754496
author Wu, Chengwei
Wang, Jiapu
Gao, Mingyang
Zhuo, Xingrui
Guo, Jipeng
Lei, Runlin
Luo, Haoran
Chen, Tianyu
Zhou, Haoyi
Pan, Shirui
Li, Zechao
author_facet Wu, Chengwei
Wang, Jiapu
Gao, Mingyang
Zhuo, Xingrui
Guo, Jipeng
Lei, Runlin
Luo, Haoran
Chen, Tianyu
Zhou, Haoyi
Pan, Shirui
Li, Zechao
contents Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks. However, Chinese LLMs face unique challenges, primarily due to the dominance of unstructured free text and the lack of structured representations in Chinese corpora. While existing benchmarks for LLMs partially assess Chinese LLMs, they are still predominantly English-centric and fail to address the unique linguistic characteristics of Chinese, lacking structured datasets essential for robust evaluation. To address these challenges, we present a Comprehensive Benchmark for Evaluating Chinese Large Language Models (CB-ECLLM) based on the newly constructed Chinese Data-Text Pair (CDTP) dataset. Specifically, CDTP comprises over 7 million aligned text pairs, each consisting of unstructured text coupled with one or more corresponding triples, alongside a total of 15 million triples spanning four critical domains. The core contributions of CDTP are threefold: (i) enriching Chinese corpora with high-quality structured information; (ii) enabling fine-grained evaluation tailored to knowledge-driven tasks; and (iii) supporting multi-task fine-tuning to assess generalization and robustness across scenarios, including Knowledge Graph Completion, Triple-to-Text generation, and Question Answering. Furthermore, we conduct rigorous evaluations through extensive experiments and ablation studies to assess the effectiveness, Supervised Fine-Tuning (SFT), and robustness of the benchmark. To support reproducible research, we offer an open-source codebase and outline potential directions for future investigations based on our insights.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
Wu, Chengwei
Wang, Jiapu
Gao, Mingyang
Zhuo, Xingrui
Guo, Jipeng
Lei, Runlin
Luo, Haoran
Chen, Tianyu
Zhou, Haoyi
Pan, Shirui
Li, Zechao
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks. However, Chinese LLMs face unique challenges, primarily due to the dominance of unstructured free text and the lack of structured representations in Chinese corpora. While existing benchmarks for LLMs partially assess Chinese LLMs, they are still predominantly English-centric and fail to address the unique linguistic characteristics of Chinese, lacking structured datasets essential for robust evaluation. To address these challenges, we present a Comprehensive Benchmark for Evaluating Chinese Large Language Models (CB-ECLLM) based on the newly constructed Chinese Data-Text Pair (CDTP) dataset. Specifically, CDTP comprises over 7 million aligned text pairs, each consisting of unstructured text coupled with one or more corresponding triples, alongside a total of 15 million triples spanning four critical domains. The core contributions of CDTP are threefold: (i) enriching Chinese corpora with high-quality structured information; (ii) enabling fine-grained evaluation tailored to knowledge-driven tasks; and (iii) supporting multi-task fine-tuning to assess generalization and robustness across scenarios, including Knowledge Graph Completion, Triple-to-Text generation, and Question Answering. Furthermore, we conduct rigorous evaluations through extensive experiments and ablation studies to assess the effectiveness, Supervised Fine-Tuning (SFT), and robustness of the benchmark. To support reproducible research, we offer an open-source codebase and outline potential directions for future investigations based on our insights.
title CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.06039