How Far Can Unsupervised RLVR Scale LLM Training?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: He, Bingxiang, Zuo, Yuxin, Liu, Zeyuan, Zhao, Shangziqi, Fu, Zixuan, Yang, Junlin, Qian, Cheng, Zhang, Kaiyan, Fan, Yuchen, Cui, Ganqu, Chen, Xiusi, Sun, Youbang, Lv, Xingtai, Zhu, Xuekai, Sheng, Li, Li, Ran, Gao, Huan-ang, Zhang, Yuchen, Zhou, Bowen, Liu, Zhiyuan, Ding, Ning
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912956124495872
author He, Bingxiang
Zuo, Yuxin
Liu, Zeyuan
Zhao, Shangziqi
Fu, Zixuan
Yang, Junlin
Qian, Cheng
Zhang, Kaiyan
Fan, Yuchen
Cui, Ganqu
Chen, Xiusi
Sun, Youbang
Lv, Xingtai
Zhu, Xuekai
Sheng, Li
Li, Ran
Gao, Huan-ang
Zhang, Yuchen
Zhou, Bowen
Liu, Zhiyuan
Ding, Ning
author_facet He, Bingxiang
Zuo, Yuxin
Liu, Zeyuan
Zhao, Shangziqi
Fu, Zixuan
Yang, Junlin
Qian, Cheng
Zhang, Kaiyan
Fan, Yuchen
Cui, Ganqu
Chen, Xiusi
Sun, Youbang
Lv, Xingtai
Zhu, Xuekai
Sheng, Li
Li, Ran
Gao, Huan-ang
Zhang, Yuchen
Zhou, Bowen
Liu, Zhiyuan
Ding, Ning
contents Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning taxonomy, theory and extensive experiments. We first classify URLVR methods into intrinsic versus external based on reward sources, then establish a unified theoretical framework revealing that all intrinsic methods converge toward sharpening the model's initial distribution This sharpening mechanism succeeds when initial confidence aligns with correctness but fails catastrophically when misaligned. Through systematic experiments, we show intrinsic rewards consistently follow a rise-then-fall pattern across methods, with collapse timing determined by model prior rather than engineering choices. Despite these scaling limits, we find intrinsic rewards remain valuable in test-time training on small datasets, and propose Model Collapse Step to measure model prior, serving as a practical indicator for RL trainability. Finally, we explore external reward methods that ground verification in computational asymmetries, showing preliminary evidence they may escape the confidence-correctness ceiling. Our findings chart boundaries for intrinsic URLVR while motivating paths toward scalable alternatives.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08660
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Far Can Unsupervised RLVR Scale LLM Training?
He, Bingxiang
Zuo, Yuxin
Liu, Zeyuan
Zhao, Shangziqi
Fu, Zixuan
Yang, Junlin
Qian, Cheng
Zhang, Kaiyan
Fan, Yuchen
Cui, Ganqu
Chen, Xiusi
Sun, Youbang
Lv, Xingtai
Zhu, Xuekai
Sheng, Li
Li, Ran
Gao, Huan-ang
Zhang, Yuchen
Zhou, Bowen
Liu, Zhiyuan
Ding, Ning
Machine Learning
Computation and Language
Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning taxonomy, theory and extensive experiments. We first classify URLVR methods into intrinsic versus external based on reward sources, then establish a unified theoretical framework revealing that all intrinsic methods converge toward sharpening the model's initial distribution This sharpening mechanism succeeds when initial confidence aligns with correctness but fails catastrophically when misaligned. Through systematic experiments, we show intrinsic rewards consistently follow a rise-then-fall pattern across methods, with collapse timing determined by model prior rather than engineering choices. Despite these scaling limits, we find intrinsic rewards remain valuable in test-time training on small datasets, and propose Model Collapse Step to measure model prior, serving as a practical indicator for RL trainability. Finally, we explore external reward methods that ground verification in computational asymmetries, showing preliminary evidence they may escape the confidence-correctness ceiling. Our findings chart boundaries for intrinsic URLVR while motivating paths toward scalable alternatives.
title How Far Can Unsupervised RLVR Scale LLM Training?
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2603.08660