Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Bo, Kong, Deqian, Savarese, Silvio, Xiong, Caiming, Zhou, Yingbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917050177290240
author Pang, Bo
Kong, Deqian
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
author_facet Pang, Bo
Kong, Deqian
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
contents Reinforcement learning (RL) can elicit strong reasoning in large language models (LLMs), yet most open efforts focus on math and code. We propose Reasoning Curriculum, a simple two-stage curriculum that first elicits reasoning skills in pretraining-aligned domains such as math, then adapts and refines these skills across other domains via joint RL. Stage 1 performs a brief cold start and then math-only RL with verifiable rewards to develop reasoning skills. Stage 2 runs joint RL on mixed-domain data to transfer and consolidate these skills. The curriculum is minimal and backbone-agnostic, requiring no specialized reward models beyond standard verifiability checks. Evaluated on Qwen3-4B and Llama-3.1-8B over a multi-domain suite, reasoning curriculum yields consistent gains. Ablations and a cognitive-skill analysis indicate that both stages are necessary and that math-first elicitation increases cognitive behaviors important for solving complex problems. Reasoning Curriculum provides a compact, easy-to-adopt recipe for general reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
Pang, Bo
Kong, Deqian
Savarese, Silvio
Xiong, Caiming
Zhou, Yingbo
Artificial Intelligence
Computation and Language
Reinforcement learning (RL) can elicit strong reasoning in large language models (LLMs), yet most open efforts focus on math and code. We propose Reasoning Curriculum, a simple two-stage curriculum that first elicits reasoning skills in pretraining-aligned domains such as math, then adapts and refines these skills across other domains via joint RL. Stage 1 performs a brief cold start and then math-only RL with verifiable rewards to develop reasoning skills. Stage 2 runs joint RL on mixed-domain data to transfer and consolidate these skills. The curriculum is minimal and backbone-agnostic, requiring no specialized reward models beyond standard verifiability checks. Evaluated on Qwen3-4B and Llama-3.1-8B over a multi-domain suite, reasoning curriculum yields consistent gains. Ablations and a cognitive-skill analysis indicate that both stages are necessary and that math-first elicitation increases cognitive behaviors important for solving complex problems. Reasoning Curriculum provides a compact, easy-to-adopt recipe for general reasoning.
title Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.26143