DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shao, Zhihong, Wang, Peiyi, Zhu, Qihao, Xu, Runxin, Song, Junxiao, Bi, Xiao, Zhang, Haowei, Zhang, Mingchuan, Li, Y. K., Wu, Y., Guo, Daya
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914773210234880
author Shao, Zhihong
Wang, Peiyi
Zhu, Qihao
Xu, Runxin
Song, Junxiao
Bi, Xiao
Zhang, Haowei
Zhang, Mingchuan
Li, Y. K.
Wu, Y.
Guo, Daya
author_facet Shao, Zhihong
Wang, Peiyi
Zhu, Qihao
Xu, Runxin
Song, Junxiao
Bi, Xiao
Zhang, Haowei
Zhang, Mingchuan
Li, Y. K.
Wu, Y.
Guo, Daya
contents Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03300
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao, Zhihong
Wang, Peiyi
Zhu, Qihao
Xu, Runxin
Song, Junxiao
Bi, Xiao
Zhang, Haowei
Zhang, Mingchuan
Li, Y. K.
Wu, Y.
Guo, Daya
Computation and Language
Artificial Intelligence
Machine Learning
Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.
title DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.03300