CoRT: Code-integrated Reasoning within Thinking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chengpeng, Tang, Zhengyang, Li, Ziniu, Xue, Mingfeng, Bao, Keqin, Ding, Tian, Sun, Ruoyu, Wang, Benyou, Wang, Xiang, Lin, Junyang, Liu, Dayiheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909646722170880
author Li, Chengpeng
Tang, Zhengyang
Li, Ziniu
Xue, Mingfeng
Bao, Keqin
Ding, Tian
Sun, Ruoyu
Wang, Benyou
Wang, Xiang
Lin, Junyang
Liu, Dayiheng
author_facet Li, Chengpeng
Tang, Zhengyang
Li, Ziniu
Xue, Mingfeng
Bao, Keqin
Ding, Tian
Sun, Ruoyu
Wang, Benyou
Wang, Xiang
Lin, Junyang
Liu, Dayiheng
contents Large Reasoning Models (LRMs) like o1 and DeepSeek-R1 have shown remarkable progress in natural language reasoning with long chain-of-thought (CoT), yet they remain inefficient or inaccurate when handling complex mathematical operations. Addressing these limitations through computational tools (e.g., computation libraries and symbolic solvers) is promising, but it introduces a technical challenge: Code Interpreter (CI) brings external knowledge beyond the model's internal text representations, thus the direct combination is not efficient. This paper introduces CoRT, a post-training framework for teaching LRMs to leverage CI effectively and efficiently. As a first step, we address the data scarcity issue by synthesizing code-integrated reasoning data through Hint-Engineering, which strategically inserts different hints at appropriate positions to optimize LRM-CI interaction. We manually create 30 high-quality samples, upon which we post-train models ranging from 1.5B to 32B parameters, with supervised fine-tuning, rejection fine-tuning and reinforcement learning. Our experimental results demonstrate that Hint-Engineering models achieve 4\% and 8\% absolute improvements on DeepSeek-R1-Distill-Qwen-32B and DeepSeek-R1-Distill-Qwen-1.5B respectively, across five challenging mathematical reasoning datasets. Furthermore, Hint-Engineering models use about 30\% fewer tokens for the 32B model and 50\% fewer tokens for the 1.5B model compared with the natural language models. The models and code are available at https://github.com/ChengpengLi1003/CoRT.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoRT: Code-integrated Reasoning within Thinking
Li, Chengpeng
Tang, Zhengyang
Li, Ziniu
Xue, Mingfeng
Bao, Keqin
Ding, Tian
Sun, Ruoyu
Wang, Benyou
Wang, Xiang
Lin, Junyang
Liu, Dayiheng
Computation and Language
Artificial Intelligence
Machine Learning
Large Reasoning Models (LRMs) like o1 and DeepSeek-R1 have shown remarkable progress in natural language reasoning with long chain-of-thought (CoT), yet they remain inefficient or inaccurate when handling complex mathematical operations. Addressing these limitations through computational tools (e.g., computation libraries and symbolic solvers) is promising, but it introduces a technical challenge: Code Interpreter (CI) brings external knowledge beyond the model's internal text representations, thus the direct combination is not efficient. This paper introduces CoRT, a post-training framework for teaching LRMs to leverage CI effectively and efficiently. As a first step, we address the data scarcity issue by synthesizing code-integrated reasoning data through Hint-Engineering, which strategically inserts different hints at appropriate positions to optimize LRM-CI interaction. We manually create 30 high-quality samples, upon which we post-train models ranging from 1.5B to 32B parameters, with supervised fine-tuning, rejection fine-tuning and reinforcement learning. Our experimental results demonstrate that Hint-Engineering models achieve 4\% and 8\% absolute improvements on DeepSeek-R1-Distill-Qwen-32B and DeepSeek-R1-Distill-Qwen-1.5B respectively, across five challenging mathematical reasoning datasets. Furthermore, Hint-Engineering models use about 30\% fewer tokens for the 32B model and 50\% fewer tokens for the 1.5B model compared with the natural language models. The models and code are available at https://github.com/ChengpengLi1003/CoRT.
title CoRT: Code-integrated Reasoning within Thinking
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.09820