OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Shaobo, Ouyang, Xuan, Xu, Tianyi, Hu, Yuzheng, Liu, Jialin, Chen, Guo, Zhang, Tianyu, Zheng, Junhao, Yang, Kexin, Ren, Xingzhang, Liu, Dayiheng, Zhang, Linfeng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908819020316672
author Wang, Shaobo
Ouyang, Xuan
Xu, Tianyi
Hu, Yuzheng
Liu, Jialin
Chen, Guo
Zhang, Tianyu
Zheng, Junhao
Yang, Kexin
Ren, Xingzhang
Liu, Dayiheng
Zhang, Linfeng
author_facet Wang, Shaobo
Ouyang, Xuan
Xu, Tianyi
Hu, Yuzheng
Liu, Jialin
Chen, Guo
Zhang, Tianyu
Zheng, Junhao
Yang, Kexin
Ren, Xingzhang
Liu, Dayiheng
Zhang, Linfeng
contents As high-quality public text approaches exhaustion, a phenomenon known as the Data Wall, pre-training is shifting from more tokens to better tokens. However, existing methods either rely on heuristic static filters that ignore training dynamics, or use dynamic yet optimizer-agnostic criteria based on raw gradients. We propose OPUS (Optimizer-induced Projected Utility Selection), a dynamic data selection framework that defines utility in the optimizer-induced update space. OPUS scores candidates by projecting their effective updates, shaped by modern optimizers, onto a target direction derived from a stable, in-distribution proxy. To ensure scalability, we employ Ghost technique with CountSketch for computational efficiency, and Boltzmann sampling for data diversity, incurring only 4.7\% additional compute overhead. OPUS achieves remarkable results across diverse corpora, quality tiers, optimizers, and model scales. In pre-training of GPT-2 Large/XL on FineWeb and FineWeb-Edu with 30B tokens, OPUS outperforms industrial-level baselines and even full 200B-token training. Moreover, when combined with industrial-level static filters, OPUS further improves pre-training efficiency, even with lower-quality data. Furthermore, in continued pre-training of Qwen3-8B-Base on SciencePedia, OPUS achieves superior performance using only 0.5B tokens compared to full training with 3B tokens, demonstrating significant data efficiency gains in specialized domains.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05400
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
Wang, Shaobo
Ouyang, Xuan
Xu, Tianyi
Hu, Yuzheng
Liu, Jialin
Chen, Guo
Zhang, Tianyu
Zheng, Junhao
Yang, Kexin
Ren, Xingzhang
Liu, Dayiheng
Zhang, Linfeng
Computation and Language
As high-quality public text approaches exhaustion, a phenomenon known as the Data Wall, pre-training is shifting from more tokens to better tokens. However, existing methods either rely on heuristic static filters that ignore training dynamics, or use dynamic yet optimizer-agnostic criteria based on raw gradients. We propose OPUS (Optimizer-induced Projected Utility Selection), a dynamic data selection framework that defines utility in the optimizer-induced update space. OPUS scores candidates by projecting their effective updates, shaped by modern optimizers, onto a target direction derived from a stable, in-distribution proxy. To ensure scalability, we employ Ghost technique with CountSketch for computational efficiency, and Boltzmann sampling for data diversity, incurring only 4.7\% additional compute overhead. OPUS achieves remarkable results across diverse corpora, quality tiers, optimizers, and model scales. In pre-training of GPT-2 Large/XL on FineWeb and FineWeb-Edu with 30B tokens, OPUS outperforms industrial-level baselines and even full 200B-token training. Moreover, when combined with industrial-level static filters, OPUS further improves pre-training efficiency, even with lower-quality data. Furthermore, in continued pre-training of Qwen3-8B-Base on SciencePedia, OPUS achieves superior performance using only 0.5B tokens compared to full training with 3B tokens, demonstrating significant data efficiency gains in specialized domains.
title OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
topic Computation and Language
url https://arxiv.org/abs/2602.05400