Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Longze, Liu, Ziqiang, He, Wanwei, Li, Yunshui, Luo, Run, Yang, Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910460930949120
author Chen, Longze
Liu, Ziqiang
He, Wanwei
Li, Yunshui
Luo, Run
Yang, Min
author_facet Chen, Longze
Liu, Ziqiang
He, Wanwei
Li, Yunshui
Luo, Run
Yang, Min
contents Long-context modeling capabilities are important for large language models (LLMs) in various applications. However, directly training LLMs with long context windows is insufficient to enhance this capability since some training samples do not exhibit strong semantic dependencies across long contexts. In this study, we propose a data mining framework \textbf{ProLong} that can assign each training sample with a long dependency score, which can be used to rank and filter samples that are more advantageous for enhancing long-context modeling abilities in LLM training. Specifically, we first use delta perplexity scores to measure the \textit{Dependency Strength} between text segments in a given document. Then we refine this metric based on the \textit{Dependency Distance} of these segments to incorporate spatial relationships across long-contexts. Final results are calibrated with a \textit{Dependency Specificity} metric to prevent trivial dependencies introduced by repetitive patterns. Moreover, a random sampling approach is proposed to optimize the computational efficiency of ProLong. Comprehensive experiments on multiple benchmarks indicate that ProLong effectively identifies documents that carry long dependencies and LLMs trained on these documents exhibit significantly enhanced long-context modeling capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17915
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
Chen, Longze
Liu, Ziqiang
He, Wanwei
Li, Yunshui
Luo, Run
Yang, Min
Computation and Language
Long-context modeling capabilities are important for large language models (LLMs) in various applications. However, directly training LLMs with long context windows is insufficient to enhance this capability since some training samples do not exhibit strong semantic dependencies across long contexts. In this study, we propose a data mining framework \textbf{ProLong} that can assign each training sample with a long dependency score, which can be used to rank and filter samples that are more advantageous for enhancing long-context modeling abilities in LLM training. Specifically, we first use delta perplexity scores to measure the \textit{Dependency Strength} between text segments in a given document. Then we refine this metric based on the \textit{Dependency Distance} of these segments to incorporate spatial relationships across long-contexts. Final results are calibrated with a \textit{Dependency Specificity} metric to prevent trivial dependencies introduced by repetitive patterns. Moreover, a random sampling approach is proposed to optimize the computational efficiency of ProLong. Comprehensive experiments on multiple benchmarks indicate that ProLong effectively identifies documents that carry long dependencies and LLMs trained on these documents exhibit significantly enhanced long-context modeling capabilities.
title Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2405.17915