DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Qiying, Zhang, Zheng, Zhu, Ruofei, Yuan, Yufeng, Zuo, Xiaochen, Yue, Yu, Dai, Weinan, Fan, Tiantian, Liu, Gaohong, Liu, Lingjun, Liu, Xin, Lin, Haibin, Lin, Zhiqi, Ma, Bole, Sheng, Guangming, Tong, Yuxuan, Zhang, Chi, Zhang, Mofan, Zhang, Wang, Zhu, Hang, Zhu, Jinhua, Chen, Jiaze, Chen, Jiangjie, Wang, Chengyi, Yu, Hongli, Song, Yuxuan, Wei, Xiangpeng, Zhou, Hao, Liu, Jingjing, Ma, Wei-Ying, Zhang, Ya-Qin, Yan, Lin, Qiao, Mu, Wu, Yonghui, Wang, Mingxuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909616714022912
author Yu, Qiying
Zhang, Zheng
Zhu, Ruofei
Yuan, Yufeng
Zuo, Xiaochen
Yue, Yu
Dai, Weinan
Fan, Tiantian
Liu, Gaohong
Liu, Lingjun
Liu, Xin
Lin, Haibin
Lin, Zhiqi
Ma, Bole
Sheng, Guangming
Tong, Yuxuan
Zhang, Chi
Zhang, Mofan
Zhang, Wang
Zhu, Hang
Zhu, Jinhua
Chen, Jiaze
Chen, Jiangjie
Wang, Chengyi
Yu, Hongli
Song, Yuxuan
Wei, Xiangpeng
Zhou, Hao
Liu, Jingjing
Ma, Wei-Ying
Zhang, Ya-Qin
Yan, Lin
Qiao, Mu
Wu, Yonghui
Wang, Mingxuan
author_facet Yu, Qiying
Zhang, Zheng
Zhu, Ruofei
Yuan, Yufeng
Zuo, Xiaochen
Yue, Yu
Dai, Weinan
Fan, Tiantian
Liu, Gaohong
Liu, Lingjun
Liu, Xin
Lin, Haibin
Lin, Zhiqi
Ma, Bole
Sheng, Guangming
Tong, Yuxuan
Zhang, Chi
Zhang, Mofan
Zhang, Wang
Zhu, Hang
Zhu, Jinhua
Chen, Jiaze
Chen, Jiangjie
Wang, Chengyi
Yu, Hongli
Song, Yuxuan
Wei, Xiangpeng
Zhou, Hao
Liu, Jingjing
Ma, Wei-Ying
Zhang, Ya-Qin
Yan, Lin
Qiao, Mu
Wu, Yonghui
Wang, Mingxuan
contents Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the $\textbf{D}$ecoupled Clip and $\textbf{D}$ynamic s$\textbf{A}$mpling $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{DAPO}$) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Yu, Qiying
Zhang, Zheng
Zhu, Ruofei
Yuan, Yufeng
Zuo, Xiaochen
Yue, Yu
Dai, Weinan
Fan, Tiantian
Liu, Gaohong
Liu, Lingjun
Liu, Xin
Lin, Haibin
Lin, Zhiqi
Ma, Bole
Sheng, Guangming
Tong, Yuxuan
Zhang, Chi
Zhang, Mofan
Zhang, Wang
Zhu, Hang
Zhu, Jinhua
Chen, Jiaze
Chen, Jiangjie
Wang, Chengyi
Yu, Hongli
Song, Yuxuan
Wei, Xiangpeng
Zhou, Hao
Liu, Jingjing
Ma, Wei-Ying
Zhang, Ya-Qin
Yan, Lin
Qiao, Mu
Wu, Yonghui
Wang, Mingxuan
Machine Learning
Computation and Language
Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the $\textbf{D}$ecoupled Clip and $\textbf{D}$ynamic s$\textbf{A}$mpling $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{DAPO}$) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL.
title DAPO: An Open-Source LLM Reinforcement Learning System at Scale
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2503.14476