Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cen, Zhepeng, Chen, Haolin, Wang, Shiyu, Liu, Zuxin, Liu, Zhiwei, Qiu, Jielin, Zhao, Ding, Savarese, Silvio, Xiong, Caiming, Wang, Huan, Yao, Weiran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911579116666880
author Cen, Zhepeng
Chen, Haolin
Wang, Shiyu
Liu, Zuxin
Liu, Zhiwei
Qiu, Jielin
Zhao, Ding
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Yao, Weiran
author_facet Cen, Zhepeng
Chen, Haolin
Wang, Shiyu
Liu, Zuxin
Liu, Zhiwei
Qiu, Jielin
Zhao, Ding
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Yao, Weiran
contents Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more data-efficient solution capable of bridging this gap, yet its application has been constrained by a critical data bottleneck: existing RL datasets are orders of magnitude smaller and less diverse than web-scale pre-training corpora. To address this, we introduce the Webscale-RL pipeline, a scalable data engine that systematically converts large-scale pre-training documents into millions of diverse, verifiable question-answer pairs for RL. Using this pipeline, we construct the Webscale-RL dataset, containing 1.2 million examples across more than 9 domains. Our experiments show that the model trained on this dataset significantly outperforms continual pretraining and strong data refinement baselines across a suite of benchmarks. Notably, RL training with our dataset proves substantially more efficient, achieving the performance of continual pre-training with up to 100$\times$ fewer tokens. Our work presents a viable path toward scaling RL to pre-training levels, enabling more capable and efficient language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06499
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
Cen, Zhepeng
Chen, Haolin
Wang, Shiyu
Liu, Zuxin
Liu, Zhiwei
Qiu, Jielin
Zhao, Ding
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Yao, Weiran
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more data-efficient solution capable of bridging this gap, yet its application has been constrained by a critical data bottleneck: existing RL datasets are orders of magnitude smaller and less diverse than web-scale pre-training corpora. To address this, we introduce the Webscale-RL pipeline, a scalable data engine that systematically converts large-scale pre-training documents into millions of diverse, verifiable question-answer pairs for RL. Using this pipeline, we construct the Webscale-RL dataset, containing 1.2 million examples across more than 9 domains. Our experiments show that the model trained on this dataset significantly outperforms continual pretraining and strong data refinement baselines across a suite of benchmarks. Notably, RL training with our dataset proves substantially more efficient, achieving the performance of continual pre-training with up to 100$\times$ fewer tokens. Our work presents a viable path toward scaling RL to pre-training levels, enabling more capable and efficient language models.
title Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.06499