Understanding R1-Zero-Like Training: A Critical Perspective

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zichen, Chen, Changyu, Li, Wenjun, Qi, Penghui, Pang, Tianyu, Du, Chao, Lee, Wee Sun, Lin, Min
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915533190856704
author Liu, Zichen
Chen, Changyu
Li, Wenjun
Qi, Penghui
Pang, Tianyu
Du, Chao
Lee, Wee Sun
Lin, Min
author_facet Liu, Zichen
Chen, Changyu
Li, Wenjun
Qi, Penghui
Pang, Tianyu
Du, Chao
Lee, Wee Sun
Lin, Min
contents DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervised fine-tuning. In this work, we critically examine R1-Zero-like training by analyzing its two core components: base models and RL. We investigate a wide range of base models, including DeepSeek-V3-Base, to understand how pretraining characteristics influence RL performance. Our analysis reveals that DeepSeek-V3-Base already exhibit ''Aha moment'', while Qwen2.5 base models demonstrate strong reasoning capabilities even without prompt templates, suggesting potential pretraining biases. Additionally, we identify an optimization bias in Group Relative Policy Optimization (GRPO), which artificially increases response length (especially for incorrect outputs) during training. To address this, we introduce Dr. GRPO, an unbiased optimization method that improves token efficiency while maintaining reasoning performance. Leveraging these insights, we present a minimalist R1-Zero recipe that achieves 43.3% accuracy on AIME 2024 with a 7B base model, establishing a new state-of-the-art. Our code is available at https://github.com/sail-sg/understand-r1-zero.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20783
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding R1-Zero-Like Training: A Critical Perspective
Liu, Zichen
Chen, Changyu
Li, Wenjun
Qi, Penghui
Pang, Tianyu
Du, Chao
Lee, Wee Sun
Lin, Min
Machine Learning
Artificial Intelligence
Computation and Language
DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervised fine-tuning. In this work, we critically examine R1-Zero-like training by analyzing its two core components: base models and RL. We investigate a wide range of base models, including DeepSeek-V3-Base, to understand how pretraining characteristics influence RL performance. Our analysis reveals that DeepSeek-V3-Base already exhibit ''Aha moment'', while Qwen2.5 base models demonstrate strong reasoning capabilities even without prompt templates, suggesting potential pretraining biases. Additionally, we identify an optimization bias in Group Relative Policy Optimization (GRPO), which artificially increases response length (especially for incorrect outputs) during training. To address this, we introduce Dr. GRPO, an unbiased optimization method that improves token efficiency while maintaining reasoning performance. Leveraging these insights, we present a minimalist R1-Zero recipe that achieves 43.3% accuracy on AIME 2024 with a 7B base model, establishing a new state-of-the-art. Our code is available at https://github.com/sail-sg/understand-r1-zero.
title Understanding R1-Zero-Like Training: A Critical Perspective
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.20783