Enhancing LLMs via High-Knowledge Data Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Feiyu, Zhang, Xuemiao, Wang, Sirui, Que, Haoran, Liu, Yuqi, Rong, Wenge, Cai, Xunliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916770208546816
author Duan, Feiyu
Zhang, Xuemiao
Wang, Sirui
Que, Haoran
Liu, Yuqi
Rong, Wenge
Cai, Xunliang
author_facet Duan, Feiyu
Zhang, Xuemiao
Wang, Sirui
Que, Haoran
Liu, Yuqi
Rong, Wenge
Cai, Xunliang
contents The performance of Large Language Models (LLMs) is intrinsically linked to the quality of its training data. Although several studies have proposed methods for high-quality data selection, they do not consider the importance of knowledge richness in text corpora. In this paper, we propose a novel and gradient-free High-Knowledge Scorer (HKS) to select high-quality data from the dimension of knowledge, to alleviate the problem of knowledge scarcity in the pre-trained corpus. We propose a comprehensive multi-domain knowledge element pool and introduce knowledge density and coverage as metrics to assess the knowledge content of the text. Based on this, we propose a comprehensive knowledge scorer to select data with intensive knowledge, which can also be utilized for domain-specific high-knowledge data selection by restricting knowledge elements to the specific domain. We train models on a high-knowledge bilingual dataset, and experimental results demonstrate that our scorer improves the model's performance in knowledge-intensive and general comprehension tasks, and is effective in enhancing both the generic and domain-specific capabilities of the model.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14070
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing LLMs via High-Knowledge Data Selection
Duan, Feiyu
Zhang, Xuemiao
Wang, Sirui
Que, Haoran
Liu, Yuqi
Rong, Wenge
Cai, Xunliang
Computation and Language
The performance of Large Language Models (LLMs) is intrinsically linked to the quality of its training data. Although several studies have proposed methods for high-quality data selection, they do not consider the importance of knowledge richness in text corpora. In this paper, we propose a novel and gradient-free High-Knowledge Scorer (HKS) to select high-quality data from the dimension of knowledge, to alleviate the problem of knowledge scarcity in the pre-trained corpus. We propose a comprehensive multi-domain knowledge element pool and introduce knowledge density and coverage as metrics to assess the knowledge content of the text. Based on this, we propose a comprehensive knowledge scorer to select data with intensive knowledge, which can also be utilized for domain-specific high-knowledge data selection by restricting knowledge elements to the specific domain. We train models on a high-knowledge bilingual dataset, and experimental results demonstrate that our scorer improves the model's performance in knowledge-intensive and general comprehension tasks, and is effective in enhancing both the generic and domain-specific capabilities of the model.
title Enhancing LLMs via High-Knowledge Data Selection
topic Computation and Language
url https://arxiv.org/abs/2505.14070