Skywork Open Reasoner 1 Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Jujie, Liu, Jiacai, Liu, Chris Yuhao, Yan, Rui, Wang, Chaojie, Cheng, Peng, Zhang, Xiaoyu, Zhang, Fuxiang, Xu, Jiacheng, Shen, Wei, Li, Siyuan, Zeng, Liang, Wei, Tianwen, Cheng, Cheng, An, Bo, Liu, Yang, Zhou, Yahui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913865817653248
author He, Jujie
Liu, Jiacai
Liu, Chris Yuhao
Yan, Rui
Wang, Chaojie
Cheng, Peng
Zhang, Xiaoyu
Zhang, Fuxiang
Xu, Jiacheng
Shen, Wei
Li, Siyuan
Zeng, Liang
Wei, Tianwen
Cheng, Cheng
An, Bo
Liu, Yang
Zhou, Yahui
author_facet He, Jujie
Liu, Jiacai
Liu, Chris Yuhao
Yan, Rui
Wang, Chaojie
Cheng, Peng
Zhang, Xiaoyu
Zhang, Fuxiang
Xu, Jiacheng
Shen, Wei
Li, Siyuan
Zeng, Liang
Wei, Tianwen
Cheng, Cheng
An, Bo
Liu, Yang
Zhou, Yahui
contents The success of DeepSeek-R1 underscores the significant role of reinforcement learning (RL) in enhancing the reasoning capabilities of large language models (LLMs). In this work, we present Skywork-OR1, an effective and scalable RL implementation for long Chain-of-Thought (CoT) models. Building on the DeepSeek-R1-Distill model series, our RL approach achieves notable performance gains, increasing average accuracy across AIME24, AIME25, and LiveCodeBench from 57.8% to 72.8% (+15.0%) for the 32B model and from 43.6% to 57.5% (+13.9%) for the 7B model. Our Skywork-OR1-32B model surpasses both DeepSeek-R1 and Qwen3-32B on the AIME24 and AIME25 benchmarks, while achieving comparable results on LiveCodeBench. The Skywork-OR1-7B and Skywork-OR1-Math-7B models demonstrate competitive reasoning capabilities among models of similar size. We perform comprehensive ablation studies on the core components of our training pipeline to validate their effectiveness. Additionally, we thoroughly investigate the phenomenon of entropy collapse, identify key factors affecting entropy dynamics, and demonstrate that mitigating premature entropy collapse is critical for improved test performance. To support community research, we fully open-source our model weights, training code, and training datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Skywork Open Reasoner 1 Technical Report
He, Jujie
Liu, Jiacai
Liu, Chris Yuhao
Yan, Rui
Wang, Chaojie
Cheng, Peng
Zhang, Xiaoyu
Zhang, Fuxiang
Xu, Jiacheng
Shen, Wei
Li, Siyuan
Zeng, Liang
Wei, Tianwen
Cheng, Cheng
An, Bo
Liu, Yang
Zhou, Yahui
Machine Learning
Artificial Intelligence
Computation and Language
The success of DeepSeek-R1 underscores the significant role of reinforcement learning (RL) in enhancing the reasoning capabilities of large language models (LLMs). In this work, we present Skywork-OR1, an effective and scalable RL implementation for long Chain-of-Thought (CoT) models. Building on the DeepSeek-R1-Distill model series, our RL approach achieves notable performance gains, increasing average accuracy across AIME24, AIME25, and LiveCodeBench from 57.8% to 72.8% (+15.0%) for the 32B model and from 43.6% to 57.5% (+13.9%) for the 7B model. Our Skywork-OR1-32B model surpasses both DeepSeek-R1 and Qwen3-32B on the AIME24 and AIME25 benchmarks, while achieving comparable results on LiveCodeBench. The Skywork-OR1-7B and Skywork-OR1-Math-7B models demonstrate competitive reasoning capabilities among models of similar size. We perform comprehensive ablation studies on the core components of our training pipeline to validate their effectiveness. Additionally, we thoroughly investigate the phenomenon of entropy collapse, identify key factors affecting entropy dynamics, and demonstrate that mitigating premature entropy collapse is critical for improved test performance. To support community research, we fully open-source our model weights, training code, and training datasets.
title Skywork Open Reasoner 1 Technical Report
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.22312