JustRL: Scaling a 1.5B LLM with a Simple RL Recipe

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Bingxiang, Qu, Zekai, Liu, Zeyuan, Chen, Yinghao, Zuo, Yuxin, Qian, Cheng, Zhang, Kaiyan, Chen, Weize, Xiao, Chaojun, Cui, Ganqu, Ding, Ning, Liu, Zhiyuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909968963207168
author He, Bingxiang
Qu, Zekai
Liu, Zeyuan
Chen, Yinghao
Zuo, Yuxin
Qian, Cheng
Zhang, Kaiyan
Chen, Weize
Xiao, Chaojun
Cui, Ganqu
Ding, Ning
Liu, Zhiyuan
author_facet He, Bingxiang
Qu, Zekai
Liu, Zeyuan
Chen, Yinghao
Zuo, Yuxin
Qian, Cheng
Zhang, Kaiyan
Chen, Weize
Xiao, Chaojun
Cui, Ganqu
Ding, Ning
Liu, Zhiyuan
contents Recent advances in reinforcement learning for large language models have converged on increasing complexity: multi-stage training pipelines, dynamic hyperparameter schedules, and curriculum learning strategies. This raises a fundamental question: \textbf{Is this complexity necessary?} We present \textbf{JustRL}, a minimal approach using single-stage training with fixed hyperparameters that achieves state-of-the-art performance on two 1.5B reasoning models (54.9\% and 64.3\% average accuracy across nine mathematical benchmarks) while using 2$\times$ less compute than sophisticated approaches. The same hyperparameters transfer across both models without tuning, and training exhibits smooth, monotonic improvement over 4,000+ steps without the collapses or plateaus that typically motivate interventions. Critically, ablations reveal that adding ``standard tricks'' like explicit length penalties and robust verifiers may degrade performance by collapsing exploration. These results suggest that the field may be adding complexity to solve problems that disappear with a stable, scaled-up baseline. We release our models and code to establish a simple, validated baseline for the community.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
He, Bingxiang
Qu, Zekai
Liu, Zeyuan
Chen, Yinghao
Zuo, Yuxin
Qian, Cheng
Zhang, Kaiyan
Chen, Weize
Xiao, Chaojun
Cui, Ganqu
Ding, Ning
Liu, Zhiyuan
Computation and Language
Recent advances in reinforcement learning for large language models have converged on increasing complexity: multi-stage training pipelines, dynamic hyperparameter schedules, and curriculum learning strategies. This raises a fundamental question: \textbf{Is this complexity necessary?} We present \textbf{JustRL}, a minimal approach using single-stage training with fixed hyperparameters that achieves state-of-the-art performance on two 1.5B reasoning models (54.9\% and 64.3\% average accuracy across nine mathematical benchmarks) while using 2$\times$ less compute than sophisticated approaches. The same hyperparameters transfer across both models without tuning, and training exhibits smooth, monotonic improvement over 4,000+ steps without the collapses or plateaus that typically motivate interventions. Critically, ablations reveal that adding ``standard tricks'' like explicit length penalties and robust verifiers may degrade performance by collapsing exploration. These results suggest that the field may be adding complexity to solve problems that disappear with a stable, scaled-up baseline. We release our models and code to establish a simple, validated baseline for the community.
title JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
topic Computation and Language
url https://arxiv.org/abs/2512.16649