Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Yujun, Liang, Zhenwen, Liu, Haolin, Yu, Wenhao, Panaganti, Kishan, Song, Linfeng, Yu, Dian, Zhang, Xiangliang, Mi, Haitao, Yu, Dong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917278547705856
author Zhou, Yujun
Liang, Zhenwen
Liu, Haolin
Yu, Wenhao
Panaganti, Kishan
Song, Linfeng
Yu, Dian
Zhang, Xiangliang
Mi, Haitao
Yu, Dong
author_facet Zhou, Yujun
Liang, Zhenwen
Liu, Haolin
Yu, Wenhao
Panaganti, Kishan
Song, Linfeng
Yu, Dian
Zhang, Xiangliang
Mi, Haitao
Yu, Dong
contents Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve without labels or external judges. Existing self-improvement approaches primarily rely on self-confirmation signals (e.g., confidence, entropy, or consistency) to generate rewards. This reliance drives models toward over-confident, majority-favored solutions, causing an entropy collapse that degrades pass@n and reasoning complexity. To address this, we propose EVOL-RL, a label-free framework that mirrors the evolutionary principle of balancing selection with variation. Concretely, EVOL-RL retains the majority-voted answer as an anchor for stability, but adds a novelty-aware reward that scores each sampled solution by how different its reasoning is from other concurrently generated responses. This majority-for-stability + novelty-for-exploration rule mirrors the variation-selection principle: selection prevents drift, while novelty prevents collapse. Evaluation results show that EVOL-RL consistently outperforms the majority-only baseline; e.g., training on label-free AIME24 lifts Qwen3-4B-Base AIME25 pass@1 from baseline's 4.6% to 16.4%, and pass@16 from 18.5% to 37.9%. EVOL-RL not only prevents in-domain diversity collapse but also improves out-of-domain generalization (from math reasoning to broader tasks, e.g., MMLU-Pro and BBEH). The code is available at: https://github.com/YujunZhou/EVOL-RL.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15194
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
Zhou, Yujun
Liang, Zhenwen
Liu, Haolin
Yu, Wenhao
Panaganti, Kishan
Song, Linfeng
Yu, Dian
Zhang, Xiangliang
Mi, Haitao
Yu, Dong
Machine Learning
Computation and Language
Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR), yet real-world deployment demands models that can self-improve without labels or external judges. Existing self-improvement approaches primarily rely on self-confirmation signals (e.g., confidence, entropy, or consistency) to generate rewards. This reliance drives models toward over-confident, majority-favored solutions, causing an entropy collapse that degrades pass@n and reasoning complexity. To address this, we propose EVOL-RL, a label-free framework that mirrors the evolutionary principle of balancing selection with variation. Concretely, EVOL-RL retains the majority-voted answer as an anchor for stability, but adds a novelty-aware reward that scores each sampled solution by how different its reasoning is from other concurrently generated responses. This majority-for-stability + novelty-for-exploration rule mirrors the variation-selection principle: selection prevents drift, while novelty prevents collapse. Evaluation results show that EVOL-RL consistently outperforms the majority-only baseline; e.g., training on label-free AIME24 lifts Qwen3-4B-Base AIME25 pass@1 from baseline's 4.6% to 16.4%, and pass@16 from 18.5% to 37.9%. EVOL-RL not only prevents in-domain diversity collapse but also improves out-of-domain generalization (from math reasoning to broader tasks, e.g., MMLU-Pro and BBEH). The code is available at: https://github.com/YujunZhou/EVOL-RL.
title Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.15194