Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xuemiao, Ren, Can, Tu, Chengying, Weng, Rongxiang, Wang, Shuo, Yan, Hongfei, Wang, Jingang, Cai, Xunliang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911438989164544
author Zhang, Xuemiao
Ren, Can
Tu, Chengying
Weng, Rongxiang
Wang, Shuo
Yan, Hongfei
Wang, Jingang
Cai, Xunliang
author_facet Zhang, Xuemiao
Ren, Can
Tu, Chengying
Weng, Rongxiang
Wang, Shuo
Yan, Hongfei
Wang, Jingang
Cai, Xunliang
contents Recent progress in large reasoning models for challenging mathematical reasoning has been driven by reinforcement learning (RL). Incorporating long chain-of-thought (CoT) data during mid-training has also been shown to substantially improve reasoning depth. However, current approaches often utilize CoT data indiscriminately, leaving open the critical question of which data types most effectively enhance model reasoning capabilities. In this paper, we define the foundation model's reasoning potential for the first time as the inverse of the number of independent attempts required to correctly answer the question, which is strongly correlated with the final model performance. We then propose utilizing diverse data enriched with high-value reasoning patterns to expand the reasoning potential. Specifically, we abstract atomic reasoning patterns from CoT sequences, characterized by commonality and inductive capabilities, and use them to construct a core reference set enriched with valuable reasoning patterns. Furthermore, we propose a dual-granularity algorithm involving chains of reasoning patterns and token entropy, efficiently selecting high-value CoT data (CoTP) from the data pool that aligns with the core set, thereby training models to master reasoning effectively. Only 10B-token CoTP data enables the 85A6B Mixture-of-Experts (MoE) model to improve by 9.58% on the challenging AIME 2024 and 2025, and to raise the upper bound of downstream RL performance by 7.81%.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21124
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
Zhang, Xuemiao
Ren, Can
Tu, Chengying
Weng, Rongxiang
Wang, Shuo
Yan, Hongfei
Wang, Jingang
Cai, Xunliang
Artificial Intelligence
Computation and Language
Recent progress in large reasoning models for challenging mathematical reasoning has been driven by reinforcement learning (RL). Incorporating long chain-of-thought (CoT) data during mid-training has also been shown to substantially improve reasoning depth. However, current approaches often utilize CoT data indiscriminately, leaving open the critical question of which data types most effectively enhance model reasoning capabilities. In this paper, we define the foundation model's reasoning potential for the first time as the inverse of the number of independent attempts required to correctly answer the question, which is strongly correlated with the final model performance. We then propose utilizing diverse data enriched with high-value reasoning patterns to expand the reasoning potential. Specifically, we abstract atomic reasoning patterns from CoT sequences, characterized by commonality and inductive capabilities, and use them to construct a core reference set enriched with valuable reasoning patterns. Furthermore, we propose a dual-granularity algorithm involving chains of reasoning patterns and token entropy, efficiently selecting high-value CoT data (CoTP) from the data pool that aligns with the core set, thereby training models to master reasoning effectively. Only 10B-token CoTP data enables the 85A6B Mixture-of-Experts (MoE) model to improve by 9.58% on the challenging AIME 2024 and 2025, and to raise the upper bound of downstream RL performance by 7.81%.
title Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.21124