Towards Effective and Efficient Continual Pre-training of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911969071595520 |
|---|---|
| author | Chen, Jie Chen, Zhipeng Wang, Jiapeng Zhou, Kun Zhu, Yutao Jiang, Jinhao Min, Yingqian Zhao, Wayne Xin Dou, Zhicheng Mao, Jiaxin Lin, Yankai Song, Ruihua Xu, Jun Chen, Xu Yan, Rui Wei, Zhewei Hu, Di Huang, Wenbing Wen, Ji-Rong |
| author_facet | Chen, Jie Chen, Zhipeng Wang, Jiapeng Zhou, Kun Zhu, Yutao Jiang, Jinhao Min, Yingqian Zhao, Wayne Xin Dou, Zhicheng Mao, Jiaxin Lin, Yankai Song, Ruihua Xu, Jun Chen, Xu Yan, Rui Wei, Zhewei Hu, Di Huang, Wenbing Wen, Ji-Rong |
| contents | Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents a technical report for continually pre-training Llama-3 (8B), which significantly enhances the Chinese language ability and scientific reasoning ability of the backbone model. To enhance the new abilities while retaining the original abilities, we design specific data mixture and curriculum strategies by utilizing existing datasets and synthesizing high-quality datasets. Specifically, we synthesize multidisciplinary scientific question and answer (QA) pairs based on related web pages, and subsequently incorporate these synthetic data to improve the scientific reasoning ability of Llama-3. We refer to the model after CPT as Llama-3-SynE (Synthetic data Enhanced Llama-3). We also present the tuning experiments with a relatively small model -- TinyLlama, and employ the derived findings to train the backbone model. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of the backbone models, including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval), without hurting the original capacities. Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_18743 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards Effective and Efficient Continual Pre-training of Large Language Models Chen, Jie Chen, Zhipeng Wang, Jiapeng Zhou, Kun Zhu, Yutao Jiang, Jinhao Min, Yingqian Zhao, Wayne Xin Dou, Zhicheng Mao, Jiaxin Lin, Yankai Song, Ruihua Xu, Jun Chen, Xu Yan, Rui Wei, Zhewei Hu, Di Huang, Wenbing Wen, Ji-Rong Computation and Language 68T50 I.2.7 Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents a technical report for continually pre-training Llama-3 (8B), which significantly enhances the Chinese language ability and scientific reasoning ability of the backbone model. To enhance the new abilities while retaining the original abilities, we design specific data mixture and curriculum strategies by utilizing existing datasets and synthesizing high-quality datasets. Specifically, we synthesize multidisciplinary scientific question and answer (QA) pairs based on related web pages, and subsequently incorporate these synthetic data to improve the scientific reasoning ability of Llama-3. We refer to the model after CPT as Llama-3-SynE (Synthetic data Enhanced Llama-3). We also present the tuning experiments with a relatively small model -- TinyLlama, and employ the derived findings to train the backbone model. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of the backbone models, including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval), without hurting the original capacities. Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE. |
| title | Towards Effective and Efficient Continual Pre-training of Large Language Models |
| topic | Computation and Language 68T50 I.2.7 |
| url | https://arxiv.org/abs/2407.18743 |