DIDS: Domain Impact-aware Data Sampling for Large Language Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Weijie, Zhang, Jipeng, Wu, Yaguang, Fang, Jingzhi, Zhang, Ruiyuan, Xu, Jiajie, Zhu, Jia, Chen, Hao, Zhao, Yao, Han, Sirui, Zhou, Xiaofang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918128763535360
author Shi, Weijie
Zhang, Jipeng
Wu, Yaguang
Fang, Jingzhi
Zhang, Ruiyuan
Xu, Jiajie
Zhu, Jia
Chen, Hao
Zhao, Yao
Han, Sirui
Zhou, Xiaofang
author_facet Shi, Weijie
Zhang, Jipeng
Wu, Yaguang
Fang, Jingzhi
Zhang, Ruiyuan
Xu, Jiajie
Zhu, Jia
Chen, Hao
Zhao, Yao
Han, Sirui
Zhou, Xiaofang
contents Large language models (LLMs) are commonly trained on multi-domain datasets, where domain sampling strategies significantly impact model performance due to varying domain importance across downstream tasks. Existing approaches for optimizing domain-level sampling strategies struggle with maintaining intra-domain consistency and accurately measuring domain impact. In this paper, we present Domain Impact-aware Data Sampling (DIDS). To ensure intra-domain consistency, a gradient clustering algorithm is proposed to group training data based on their learning effects, where a proxy language model and dimensionality reduction are employed to reduce computational overhead. To accurately measure domain impact, we develop a Fisher Information Matrix (FIM) guided metric that quantifies how domain-specific parameter updates affect the model's output distributions on downstream tasks, with theoretical guarantees. Furthermore, to determine optimal sampling ratios, DIDS combines both the FIM-guided domain impact assessment and loss learning trajectories that indicate domain-specific potential, while accounting for diminishing marginal returns. Extensive experiments demonstrate that DIDS achieves 3.4% higher average performance while maintaining comparable training efficiency. The code is available at https://github.com/shiweijiezero/DIDS.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DIDS: Domain Impact-aware Data Sampling for Large Language Model Training
Shi, Weijie
Zhang, Jipeng
Wu, Yaguang
Fang, Jingzhi
Zhang, Ruiyuan
Xu, Jiajie
Zhu, Jia
Chen, Hao
Zhao, Yao
Han, Sirui
Zhou, Xiaofang
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) are commonly trained on multi-domain datasets, where domain sampling strategies significantly impact model performance due to varying domain importance across downstream tasks. Existing approaches for optimizing domain-level sampling strategies struggle with maintaining intra-domain consistency and accurately measuring domain impact. In this paper, we present Domain Impact-aware Data Sampling (DIDS). To ensure intra-domain consistency, a gradient clustering algorithm is proposed to group training data based on their learning effects, where a proxy language model and dimensionality reduction are employed to reduce computational overhead. To accurately measure domain impact, we develop a Fisher Information Matrix (FIM) guided metric that quantifies how domain-specific parameter updates affect the model's output distributions on downstream tasks, with theoretical guarantees. Furthermore, to determine optimal sampling ratios, DIDS combines both the FIM-guided domain impact assessment and loss learning trajectories that indicate domain-specific potential, while accounting for diminishing marginal returns. Extensive experiments demonstrate that DIDS achieves 3.4% higher average performance while maintaining comparable training efficiency. The code is available at https://github.com/shiweijiezero/DIDS.
title DIDS: Domain Impact-aware Data Sampling for Large Language Model Training
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.13227