Data Engineering for Scaling Language Models to 128K Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Yao, Panda, Rameswar, Niu, Xinyao, Yue, Xiang, Hajishirzi, Hannaneh, Kim, Yoon, Peng, Hao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913235249135616
author Fu, Yao
Panda, Rameswar
Niu, Xinyao
Yue, Xiang
Hajishirzi, Hannaneh
Kim, Yoon
Peng, Hao
author_facet Fu, Yao
Panda, Rameswar
Niu, Xinyao
Yue, Xiang
Hajishirzi, Hannaneh
Kim, Yoon
Peng, Hao
contents We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in particular \textit{the ability to utilize information at arbitrary input locations}, is a capability that is mostly already acquired through large-scale pretraining, and that this capability can be readily extended to contexts substantially longer than seen during training~(e.g., 4K to 128K) through lightweight continual pretraining on appropriate data mixture. We investigate the \textit{quantity} and \textit{quality} of the data for continual pretraining: (1) for quantity, we show that 500 million to 5 billion tokens are enough to enable the model to retrieve information anywhere within the 128K context; (2) for quality, our results equally emphasize \textit{domain balance} and \textit{length upsampling}. Concretely, we find that naively upsampling longer data on certain domains like books, a common practice of existing work, gives suboptimal performance, and that a balanced domain mixture is important. We demonstrate that continual pretraining of the full model on 1B-5B tokens of such data is an effective and affordable strategy for scaling the context length of language models to 128K. Our recipe outperforms strong open-source long-context models and closes the gap to frontier models like GPT-4 128K.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10171
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Engineering for Scaling Language Models to 128K Context
Fu, Yao
Panda, Rameswar
Niu, Xinyao
Yue, Xiang
Hajishirzi, Hannaneh
Kim, Yoon
Peng, Hao
Computation and Language
Artificial Intelligence
We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering. We hypothesize that long context modeling, in particular \textit{the ability to utilize information at arbitrary input locations}, is a capability that is mostly already acquired through large-scale pretraining, and that this capability can be readily extended to contexts substantially longer than seen during training~(e.g., 4K to 128K) through lightweight continual pretraining on appropriate data mixture. We investigate the \textit{quantity} and \textit{quality} of the data for continual pretraining: (1) for quantity, we show that 500 million to 5 billion tokens are enough to enable the model to retrieve information anywhere within the 128K context; (2) for quality, our results equally emphasize \textit{domain balance} and \textit{length upsampling}. Concretely, we find that naively upsampling longer data on certain domains like books, a common practice of existing work, gives suboptimal performance, and that a balanced domain mixture is important. We demonstrate that continual pretraining of the full model on 1B-5B tokens of such data is an effective and affordable strategy for scaling the context length of language models to 128K. Our recipe outperforms strong open-source long-context models and closes the gap to frontier models like GPT-4 128K.
title Data Engineering for Scaling Language Models to 128K Context
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.10171