JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hao, Yifan, Chao, Fangning, Hao, Yaqian, Cui, Zhaojun, Bai, Huan, Zhang, Haiyu, Liu, Yankai, Deng, Chao, Feng, Junlan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916864988282880
author Hao, Yifan
Chao, Fangning
Hao, Yaqian
Cui, Zhaojun
Bai, Huan
Zhang, Haiyu
Liu, Yankai
Deng, Chao
Feng, Junlan
author_facet Hao, Yifan
Chao, Fangning
Hao, Yaqian
Cui, Zhaojun
Bai, Huan
Zhang, Haiyu
Liu, Yankai
Deng, Chao
Feng, Junlan
contents Mathematical reasoning is a cornerstone of artificial general intelligence and a primary benchmark for evaluating the capabilities of Large Language Models (LLMs). While state-of-the-art models show promise, they often falter when faced with complex problems that demand deep conceptual understanding and intricate, multi-step deliberation. To address this challenge, we introduce JT-Math-8B, a series of open-source models comprising base, instruct, and thinking versions, built upon a systematic, multi-stage optimization framework. Our pre-training corpus is a high-quality, 210B-token dataset curated through a dedicated data pipeline that uses model-based validation to ensure quality and diversity. The Instruct Model is optimized for direct, concise answers through Supervised Fine-Tuning (SFT) and a GRPO-based reinforcement learning (RL) method. The Thinking Model is trained for complex problem-solving using a Long Chain-of-Thought (Long CoT) approach, combining SFT with a novel, multi-stage RL curriculum that progressively increases task difficulty and context length up to 32K tokens. JT-Math-8B achieves state-of-the-art results among open-source models of similar size, surpassing prominent models like OpenAI's O1-mini and GPT-4o , and demonstrating superior performance on competition-level mathematics.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
Hao, Yifan
Chao, Fangning
Hao, Yaqian
Cui, Zhaojun
Bai, Huan
Zhang, Haiyu
Liu, Yankai
Deng, Chao
Feng, Junlan
Computation and Language
Mathematical reasoning is a cornerstone of artificial general intelligence and a primary benchmark for evaluating the capabilities of Large Language Models (LLMs). While state-of-the-art models show promise, they often falter when faced with complex problems that demand deep conceptual understanding and intricate, multi-step deliberation. To address this challenge, we introduce JT-Math-8B, a series of open-source models comprising base, instruct, and thinking versions, built upon a systematic, multi-stage optimization framework. Our pre-training corpus is a high-quality, 210B-token dataset curated through a dedicated data pipeline that uses model-based validation to ensure quality and diversity. The Instruct Model is optimized for direct, concise answers through Supervised Fine-Tuning (SFT) and a GRPO-based reinforcement learning (RL) method. The Thinking Model is trained for complex problem-solving using a Long Chain-of-Thought (Long CoT) approach, combining SFT with a novel, multi-stage RL curriculum that progressively increases task difficulty and context length up to 32K tokens. JT-Math-8B achieves state-of-the-art results among open-source models of similar size, surpassing prominent models like OpenAI's O1-mini and GPT-4o , and demonstrating superior performance on competition-level mathematics.
title JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.19748