Towards Effective and Efficient Continual Pre-training of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jie, Chen, Zhipeng, Wang, Jiapeng, Zhou, Kun, Zhu, Yutao, Jiang, Jinhao, Min, Yingqian, Zhao, Wayne Xin, Dou, Zhicheng, Mao, Jiaxin, Lin, Yankai, Song, Ruihua, Xu, Jun, Chen, Xu, Yan, Rui, Wei, Zhewei, Hu, Di, Huang, Wenbing, Wen, Ji-Rong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911969071595520
author Chen, Jie
Chen, Zhipeng
Wang, Jiapeng
Zhou, Kun
Zhu, Yutao
Jiang, Jinhao
Min, Yingqian
Zhao, Wayne Xin
Dou, Zhicheng
Mao, Jiaxin
Lin, Yankai
Song, Ruihua
Xu, Jun
Chen, Xu
Yan, Rui
Wei, Zhewei
Hu, Di
Huang, Wenbing
Wen, Ji-Rong
author_facet Chen, Jie
Chen, Zhipeng
Wang, Jiapeng
Zhou, Kun
Zhu, Yutao
Jiang, Jinhao
Min, Yingqian
Zhao, Wayne Xin
Dou, Zhicheng
Mao, Jiaxin
Lin, Yankai
Song, Ruihua
Xu, Jun
Chen, Xu
Yan, Rui
Wei, Zhewei
Hu, Di
Huang, Wenbing
Wen, Ji-Rong
contents Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents a technical report for continually pre-training Llama-3 (8B), which significantly enhances the Chinese language ability and scientific reasoning ability of the backbone model. To enhance the new abilities while retaining the original abilities, we design specific data mixture and curriculum strategies by utilizing existing datasets and synthesizing high-quality datasets. Specifically, we synthesize multidisciplinary scientific question and answer (QA) pairs based on related web pages, and subsequently incorporate these synthetic data to improve the scientific reasoning ability of Llama-3. We refer to the model after CPT as Llama-3-SynE (Synthetic data Enhanced Llama-3). We also present the tuning experiments with a relatively small model -- TinyLlama, and employ the derived findings to train the backbone model. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of the backbone models, including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval), without hurting the original capacities. Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18743
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Effective and Efficient Continual Pre-training of Large Language Models
Chen, Jie
Chen, Zhipeng
Wang, Jiapeng
Zhou, Kun
Zhu, Yutao
Jiang, Jinhao
Min, Yingqian
Zhao, Wayne Xin
Dou, Zhicheng
Mao, Jiaxin
Lin, Yankai
Song, Ruihua
Xu, Jun
Chen, Xu
Yan, Rui
Wei, Zhewei
Hu, Di
Huang, Wenbing
Wen, Ji-Rong
Computation and Language
68T50
I.2.7
Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents a technical report for continually pre-training Llama-3 (8B), which significantly enhances the Chinese language ability and scientific reasoning ability of the backbone model. To enhance the new abilities while retaining the original abilities, we design specific data mixture and curriculum strategies by utilizing existing datasets and synthesizing high-quality datasets. Specifically, we synthesize multidisciplinary scientific question and answer (QA) pairs based on related web pages, and subsequently incorporate these synthetic data to improve the scientific reasoning ability of Llama-3. We refer to the model after CPT as Llama-3-SynE (Synthetic data Enhanced Llama-3). We also present the tuning experiments with a relatively small model -- TinyLlama, and employ the derived findings to train the backbone model. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of the backbone models, including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval), without hurting the original capacities. Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE.
title Towards Effective and Efficient Continual Pre-training of Large Language Models
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2407.18743