Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Tianyi, Yang, Ling, Wong, Zhen Hao, Sun, Fupeng, Peng, Jiahui, Zhuang, Xinlin, Zhang, Chi, Wu, Lijun, Qiu, Jiantao, Zhang, Wentao, Yuan, Binhang, He, Conghui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915332296278016
author Bai, Tianyi
Yang, Ling
Wong, Zhen Hao
Sun, Fupeng
Peng, Jiahui
Zhuang, Xinlin
Zhang, Chi
Wu, Lijun
Qiu, Jiantao
Zhang, Wentao
Yuan, Binhang
He, Conghui
author_facet Bai, Tianyi
Yang, Ling
Wong, Zhen Hao
Sun, Fupeng
Peng, Jiahui
Zhuang, Xinlin
Zhang, Chi
Wu, Lijun
Qiu, Jiantao
Zhang, Wentao
Yuan, Binhang
He, Conghui
contents Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to $10.5\%$ across multiple language model benchmarks compared to the state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08102
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
Bai, Tianyi
Yang, Ling
Wong, Zhen Hao
Sun, Fupeng
Peng, Jiahui
Zhuang, Xinlin
Zhang, Chi
Wu, Lijun
Qiu, Jiantao
Zhang, Wentao
Yuan, Binhang
He, Conghui
Computation and Language
Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to $10.5\%$ across multiple language model benchmarks compared to the state-of-the-art methods.
title Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
topic Computation and Language
url https://arxiv.org/abs/2410.08102