R-Zero: Self-Evolving Reasoning LLM from Zero Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chengsong, Yu, Wenhao, Wang, Xiaoyang, Zhang, Hongming, Li, Zongxia, Li, Ruosen, Huang, Jiaxin, Mi, Haitao, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908832776585216
author Huang, Chengsong
Yu, Wenhao
Wang, Xiaoyang
Zhang, Hongming
Li, Zongxia
Li, Ruosen
Huang, Jiaxin
Mi, Haitao
Yu, Dong
author_facet Huang, Chengsong
Yu, Wenhao
Wang, Xiaoyang
Zhang, Hongming
Li, Zongxia
Li, Ruosen
Huang, Jiaxin
Mi, Haitao
Yu, Dong
contents Self-evolving Large Language Models (LLMs) offer a scalable path toward super-intelligence by autonomously generating, refining, and learning from their own experiences. However, existing methods for training such models still rely heavily on vast human-curated tasks and labels, typically via fine-tuning or reinforcement learning, which poses a fundamental bottleneck to advancing AI systems toward capabilities beyond human intelligence. To overcome this limitation, we introduce R-Zero, a fully autonomous framework that generates its own training data from scratch. Starting from a single base LLM, R-Zero initializes two independent models with distinct roles, a Challenger and a Solver. These models are optimized separately and co-evolve through interaction: the Challenger is rewarded for proposing tasks near the edge of the Solver capability, and the Solver is rewarded for solving increasingly challenging tasks posed by the Challenger. This process yields a targeted, self-improving curriculum without any pre-existing tasks and labels. Empirically, R-Zero substantially improves reasoning capability across different backbone LLMs, e.g., boosting the Qwen3-4B-Base by +6.49 on math-reasoning benchmarks and +7.54 on general-domain reasoning benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R-Zero: Self-Evolving Reasoning LLM from Zero Data
Huang, Chengsong
Yu, Wenhao
Wang, Xiaoyang
Zhang, Hongming
Li, Zongxia
Li, Ruosen
Huang, Jiaxin
Mi, Haitao
Yu, Dong
Machine Learning
Artificial Intelligence
Computation and Language
Self-evolving Large Language Models (LLMs) offer a scalable path toward super-intelligence by autonomously generating, refining, and learning from their own experiences. However, existing methods for training such models still rely heavily on vast human-curated tasks and labels, typically via fine-tuning or reinforcement learning, which poses a fundamental bottleneck to advancing AI systems toward capabilities beyond human intelligence. To overcome this limitation, we introduce R-Zero, a fully autonomous framework that generates its own training data from scratch. Starting from a single base LLM, R-Zero initializes two independent models with distinct roles, a Challenger and a Solver. These models are optimized separately and co-evolve through interaction: the Challenger is rewarded for proposing tasks near the edge of the Solver capability, and the Solver is rewarded for solving increasingly challenging tasks posed by the Challenger. This process yields a targeted, self-improving curriculum without any pre-existing tasks and labels. Empirically, R-Zero substantially improves reasoning capability across different backbone LLMs, e.g., boosting the Qwen3-4B-Base by +6.49 on math-reasoning benchmarks and +7.54 on general-domain reasoning benchmarks.
title R-Zero: Self-Evolving Reasoning LLM from Zero Data
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.05004