CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yutong, Gao, Jiandong, Wu, Ji
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909996659245056
author Chen, Yutong
Gao, Jiandong
Wu, Ji
author_facet Chen, Yutong
Gao, Jiandong
Wu, Ji
contents Training Large Reasoning Model (LRM) is usually unstable and unpredictable, especially on hard problems or weak foundation models. We found that the current post-training scaling strategy can still improve on these cases. We propose CoScale-RL, a novel scaling strategy with better data and computational efficiency. We first scale up solutions to make problems solvable. The core idea is to collect multiple solutions for each problem, rather than simply enlarging the dataset. Then, we scale up rollout computation to stabilize Reinforcement Learning. We further leverage a model merge technique called Re-distillation to sustain or even improve computational efficiency when scaling up. Our method significantly improves data and computational efficiency, with an average 3.76$\times$ accuracy improvement on four benchmarks. CoScale-RL is able to improve an LRM's ability boundary without an extensive SFT dataset. Our method provides a new scaling direction to further improve LRM's reasoning ability.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14695
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
Chen, Yutong
Gao, Jiandong
Wu, Ji
Machine Learning
Artificial Intelligence
Training Large Reasoning Model (LRM) is usually unstable and unpredictable, especially on hard problems or weak foundation models. We found that the current post-training scaling strategy can still improve on these cases. We propose CoScale-RL, a novel scaling strategy with better data and computational efficiency. We first scale up solutions to make problems solvable. The core idea is to collect multiple solutions for each problem, rather than simply enlarging the dataset. Then, we scale up rollout computation to stabilize Reinforcement Learning. We further leverage a model merge technique called Re-distillation to sustain or even improve computational efficiency when scaling up. Our method significantly improves data and computational efficiency, with an average 3.76$\times$ accuracy improvement on four benchmarks. CoScale-RL is able to improve an LRM's ability boundary without an extensive SFT dataset. Our method provides a new scaling direction to further improve LRM's reasoning ability.
title CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.14695