Compass-Thinker-7B Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Anxiang, Zhang, Haibo, Mo, Kaixiang, Zhang, Long, Liu, Shuman, Huang, Yanhui, Liu, Yawen, Sheng, Yuepeng, Huang, Yuwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909736465596416
author Zeng, Anxiang
Zhang, Haibo
Mo, Kaixiang
Zhang, Long
Liu, Shuman
Huang, Yanhui
Liu, Yawen
Sheng, Yuepeng
Huang, Yuwei
author_facet Zeng, Anxiang
Zhang, Haibo
Mo, Kaixiang
Zhang, Long
Liu, Shuman
Huang, Yanhui
Liu, Yawen
Sheng, Yuepeng
Huang, Yuwei
contents Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning is the core technology to elicit its complex reasoning. However, conducting RL experiments directly on hyperscale models involves high computational costs and resource demands, posing significant risks. We propose the Compass-Thinker-7B model, which aims to explore the potential of Reinforcement Learning with less computational resources and costs, and provides insights for further research into RL recipes for larger models. Compass-Thinker-7B is trained from an open source model through a specially designed Reinforcement Learning Pipeline. We curate a dataset of 30k verifiable mathematics problems for the Reinforcement Learning Pipeline. By configuring data and training settings with different difficulty distributions for different stages, the potential of the model is gradually released and the training efficiency is improved. Extensive evaluations show that Compass-Thinker-7B possesses exceptional reasoning potential, and achieves superior performance on mathematics compared to the same-sized RL model. Especially in the challenging AIME2024 evaluation, Compass-Thinker-7B achieves 40% accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compass-Thinker-7B Technical Report
Zeng, Anxiang
Zhang, Haibo
Mo, Kaixiang
Zhang, Long
Liu, Shuman
Huang, Yanhui
Liu, Yawen
Sheng, Yuepeng
Huang, Yuwei
Artificial Intelligence
Recent R1-Zero-like research further demonstrates that reasoning extension has given large language models (LLMs) unprecedented reasoning capabilities, and Reinforcement Learning is the core technology to elicit its complex reasoning. However, conducting RL experiments directly on hyperscale models involves high computational costs and resource demands, posing significant risks. We propose the Compass-Thinker-7B model, which aims to explore the potential of Reinforcement Learning with less computational resources and costs, and provides insights for further research into RL recipes for larger models. Compass-Thinker-7B is trained from an open source model through a specially designed Reinforcement Learning Pipeline. We curate a dataset of 30k verifiable mathematics problems for the Reinforcement Learning Pipeline. By configuring data and training settings with different difficulty distributions for different stages, the potential of the model is gradually released and the training efficiency is improved. Extensive evaluations show that Compass-Thinker-7B possesses exceptional reasoning potential, and achieves superior performance on mathematics compared to the same-sized RL model. Especially in the challenging AIME2024 evaluation, Compass-Thinker-7B achieves 40% accuracy.
title Compass-Thinker-7B Technical Report
topic Artificial Intelligence
url https://arxiv.org/abs/2508.08909