NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Wei, Qi, Siya, Wang, Xinyu, Qian, Chen, Du, Yali, He, Yulan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915476339163136
author Liu, Wei
Qi, Siya
Wang, Xinyu
Qian, Chen
Du, Yali
He, Yulan
author_facet Liu, Wei
Qi, Siya
Wang, Xinyu
Qian, Chen
Du, Yali
He, Yulan
contents Recent advances such as DeepSeek R1-Zero highlight the effectiveness of incentive training, a reinforcement learning paradigm that computes rewards solely based on the final answer part of a language model's output, thereby encouraging the generation of intermediate reasoning steps. However, these methods fundamentally rely on external verifiers, which limits their applicability to domains like mathematics and coding where such verifiers are readily available. Although reward models can serve as verifiers, they require high-quality annotated data and are costly to train. In this work, we propose NOVER, NO-VERifier Reinforcement Learning, a general reinforcement learning framework that requires only standard supervised fine-tuning data with no need for an external verifier. NOVER enables incentive training across a wide range of text-to-text tasks and outperforms the model of the same size distilled from large reasoning models such as DeepSeek R1 671B by 7.7 percent. Moreover, the flexibility of NOVER enables new possibilities for optimizing large language models, such as inverse incentive training.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
Liu, Wei
Qi, Siya
Wang, Xinyu
Qian, Chen
Du, Yali
He, Yulan
Computation and Language
Artificial Intelligence
Machine Learning
Recent advances such as DeepSeek R1-Zero highlight the effectiveness of incentive training, a reinforcement learning paradigm that computes rewards solely based on the final answer part of a language model's output, thereby encouraging the generation of intermediate reasoning steps. However, these methods fundamentally rely on external verifiers, which limits their applicability to domains like mathematics and coding where such verifiers are readily available. Although reward models can serve as verifiers, they require high-quality annotated data and are costly to train. In this work, we propose NOVER, NO-VERifier Reinforcement Learning, a general reinforcement learning framework that requires only standard supervised fine-tuning data with no need for an external verifier. NOVER enables incentive training across a wide range of text-to-text tasks and outperforms the model of the same size distilled from large reasoning models such as DeepSeek R1 671B by 7.7 percent. Moreover, the flexibility of NOVER enables new possibilities for optimizing large language models, such as inverse incentive training.
title NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.16022