Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2510.25804 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912677990760448 |
|---|---|
| author | Deng, Haoran Lin, Yingyu Lin, Zhenghao Liu, Xiao Sun, Yizhou Ma, Yi-An Gong, Yeyun |
| author_facet | Deng, Haoran Lin, Yingyu Lin, Zhenghao Liu, Xiao Sun, Yizhou Ma, Yi-An Gong, Yeyun |
| contents | Long-context language models unlock advanced capabilities in reasoning, code generation, and document summarization by leveraging dependencies across extended spans of text. However, a significant portion of readily available long-text data lacks meaningful long-distance dependencies; most spans can be predicted using only local context. Training on such data is inefficient, making careful data selection crucial. Therefore, we introduce LongFilter, a framework for curating training data tailored to long-context pretraining. LongFilter measures the information gain provided by extended context by contrasting model predictions under long-context versus short-context settings, thereby identifying samples where long-range dependencies are essential. Experiments with LLaMA-3-8B, extending its context length from 8K to 64K, show that LongFilter efficiently selects high-quality data and yields substantial improvements on benchmarks such as HELMET, LongBench, and RULER. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_25804 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data Deng, Haoran Lin, Yingyu Lin, Zhenghao Liu, Xiao Sun, Yizhou Ma, Yi-An Gong, Yeyun Computation and Language Long-context language models unlock advanced capabilities in reasoning, code generation, and document summarization by leveraging dependencies across extended spans of text. However, a significant portion of readily available long-text data lacks meaningful long-distance dependencies; most spans can be predicted using only local context. Training on such data is inefficient, making careful data selection crucial. Therefore, we introduce LongFilter, a framework for curating training data tailored to long-context pretraining. LongFilter measures the information gain provided by extended context by contrasting model predictions under long-context versus short-context settings, thereby identifying samples where long-range dependencies are essential. Experiments with LLaMA-3-8B, extending its context length from 8K to 64K, show that LongFilter efficiently selects high-quality data and yields substantial improvements on benchmarks such as HELMET, LongBench, and RULER. |
| title | Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.25804 |