JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914513809309696 |
|---|---|
| author | Chen, Xinjie Fu, Biao Wu, Jing Chen, Guoxin Liu, Xinggao Liu, Dayiheng Liao, Minpeng |
| author_facet | Chen, Xinjie Fu, Biao Wu, Jing Chen, Guoxin Liu, Xinggao Liu, Dayiheng Liao, Minpeng |
| contents | Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefully curated reward specifications. In machine-checkable domains, label-free alternatives such as majority voting or LLM-as-a-judge remove annotation cost but can introduce false positives that destabilize training. We introduce JURY-RL, a label-free RLVR framework that decouples answer proposal from reward disposal: votes from model rollouts propose a candidate answer, and a formal verifier determines whether that candidate can receive positive reward. Concretely, only rollouts matching the plurality-voted answer are rewarded when that answer is successfully verified in Lean. When verification is inconclusive, we invoke ResZero (Residual-Zero), a fallback reward that discards the unverified plurality proposal and redistributes a zero-mean, variance-preserving signal over the residual answers. This design maintains a stable optimization gradient without reinforcing unverifiable consensus. Across three backbone models trained on mathematical data, JURY-RL consistently outperforms other label-free baselines on mathematical reasoning benchmarks and transfers competitively to code generation and general benchmarks. It attains pass@1 performance comparable to supervised ground-truth training, with superior generalization demonstrated by higher pass@k and response diversity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_25419 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR Chen, Xinjie Fu, Biao Wu, Jing Chen, Guoxin Liu, Xinggao Liu, Dayiheng Liao, Minpeng Artificial Intelligence I.2.7; I.2.6 Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefully curated reward specifications. In machine-checkable domains, label-free alternatives such as majority voting or LLM-as-a-judge remove annotation cost but can introduce false positives that destabilize training. We introduce JURY-RL, a label-free RLVR framework that decouples answer proposal from reward disposal: votes from model rollouts propose a candidate answer, and a formal verifier determines whether that candidate can receive positive reward. Concretely, only rollouts matching the plurality-voted answer are rewarded when that answer is successfully verified in Lean. When verification is inconclusive, we invoke ResZero (Residual-Zero), a fallback reward that discards the unverified plurality proposal and redistributes a zero-mean, variance-preserving signal over the residual answers. This design maintains a stable optimization gradient without reinforcing unverifiable consensus. Across three backbone models trained on mathematical data, JURY-RL consistently outperforms other label-free baselines on mathematical reasoning benchmarks and transfers competitively to code generation and general benchmarks. It attains pass@1 performance comparable to supervised ground-truth training, with superior generalization demonstrated by higher pass@k and response diversity. |
| title | JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR |
| topic | Artificial Intelligence I.2.7; I.2.6 |
| url | https://arxiv.org/abs/2604.25419 |