Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915332296278016 |
|---|---|
| author | Bai, Tianyi Yang, Ling Wong, Zhen Hao Sun, Fupeng Peng, Jiahui Zhuang, Xinlin Zhang, Chi Wu, Lijun Qiu, Jiantao Zhang, Wentao Yuan, Binhang He, Conghui |
| author_facet | Bai, Tianyi Yang, Ling Wong, Zhen Hao Sun, Fupeng Peng, Jiahui Zhuang, Xinlin Zhang, Chi Wu, Lijun Qiu, Jiantao Zhang, Wentao Yuan, Binhang He, Conghui |
| contents | Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to $10.5\%$ across multiple language model benchmarks compared to the state-of-the-art methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_08102 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration Bai, Tianyi Yang, Ling Wong, Zhen Hao Sun, Fupeng Peng, Jiahui Zhuang, Xinlin Zhang, Chi Wu, Lijun Qiu, Jiantao Zhang, Wentao Yuan, Binhang He, Conghui Computation and Language Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to $10.5\%$ across multiple language model benchmarks compared to the state-of-the-art methods. |
| title | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2410.08102 |